Companies Kyutai and Mirelo have introduced MuScriptor, an innovative open model for automatic music transcription (AMT). The technology allows for the conversion of complex multi-instrument audio recordings into structured MIDI-like note streams, providing accurate separation into instrument parts.

image
image

What Happened

Developed by Kyutai and Mirelo, the MuScriptor model is capable of recognizing and separating audio into individual parts: drums, bass, keys, and vocals. The flagship version (large) is based on 1.3 billion parameters and demonstrates a significant increase in accuracy across key F1 metrics (Onsets, Frames, Offsets) compared to existing solutions such as YourMT3+.

Context

Automatic music transcription (AMT) is a critical process for digital music production. Before MuScriptor, the process of extracting accurate MIDI data from finished mixes was labor-intensive and often required manual correction, which limited automation possibilities in this field.

Why It Matters for the Industry

For the AI industry and music software developers, MuScriptor represents a qualitative leap. The open nature of the model allows it to be used for creating high-quality datasets necessary for training new generative music models. It also simplifies the integration of automatic audio deconstruction tools into modern music production workflows.

Why It Matters for Users

Musicians and composers can use MuScriptor via a web demo or a command-line interface (CLI) to quickly turn their favorite tracks into MIDI files. This opens up possibilities for deep study of compositional structures and easy editing of individual instruments in digital audio workstations (DAWs).

What Is Not Yet Known / Limitations

At this time, detailed data regarding the model's practical applicability in high-load environments is missing, including latency metrics, inference costs, and scalability for industrial production use.

Sources

Author

Look at AI, Editorial Staff