🤖 AI Summary
This study addresses the scarcity of manual annotations and the sequence inconsistencies caused by local predictions in automatic music transcription. To overcome these challenges, we propose a unified framework integrating synthetic data supervision, structured decoding, and autoregressive distillation. Specifically, MIDI-rendered audio is leveraged to generate synthetic supervision signals, alleviating the data bottleneck. We introduce the first task-specific structured decoding approach, which employs dynamic programming to ensure global coherence of the transcribed musical scores. Furthermore, knowledge distillation is utilized to transfer the capabilities of autoregressive models into more efficient architectures, balancing accuracy with inference speed. Experimental results demonstrate that the proposed method surpasses most existing systems across eight benchmarks, significantly improving both transcription accuracy and sequential consistency.
📝 Abstract
Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequences. We present SheetSage2, a unified music transcription framework that combines synthetic data, task-specific structured decoding, and autoregressive distillation. Automatically annotated MIDI, rendered into audio, provides scalable supervision across music understanding tasks. Task-specific structured decoders integrate complementary musical cues and their temporal dependencies to produce musically coherent scores. Autoregressive distillation further retains transcription accuracy without task-specific dynamic programming at inference. Across eight benchmark collections, a single SheetSage2-AR model exceeds the listed prior systems on 12 of 15 benchmark--metric pairs in our evaluation, substantially improving over SheetSage1 and surpassing task-specific models on several benchmarks. Model weights and inference code are publicly available.