Note-Level Singing Melody Transcription for Time-Aligned Musical Score Generation

📅 2025-02-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses end-to-end audio-to-time-aligned musical score transcription—simultaneously predicting note pitch, onset/offset timestamps, and precise note durations. To tackle the scarcity of duration annotations, we introduce, for the first time, explicit note duration modeling within an end-to-end note-level transcription framework. Our approach features a duration-aware tokenization scheme and a pseudo-labeling-based data augmentation strategy. The model is trained via multi-objective joint optimization, and we propose novel evaluation metrics that jointly assess temporal precision and duration consistency. Experiments demonstrate state-of-the-art performance across multiple benchmarks. Visual analysis further confirms the method’s high accuracy and robustness in modeling note durations, particularly in challenging polyphonic vocal recordings.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Diffusion Models for VisionNatural Language Processing: Code Generation / Program Synthesis from Natural Language

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingGraph Algorithms and Modeling for the Web: Algorithms and analysis for heterogeneous, signed, attributed, multi-relational, temporal, higher-order, and annotated Web-related graphs
📝 Abstract
Automatic music transcription converts audio recordings into symbolic representations, facilitating music analysis, retrieval, and generation. A musical note is characterized by pitch, onset, and offset in an audio domain, whereas it is defined in terms of pitch and note value in a musical score domain. A time-aligned score, derived from timing information along with pitch and note value, allows matching a part of the score with the corresponding part of the music audio, enabling various applications. In this paper, we consider an extended version of the traditional note-level transcription task that recognizes onset, offset, and pitch, through including extraction of additional note value to generate a time-aligned score from an audio input. To address this new challenge, we propose an end-to-end framework that integrates recognition of the note value, pitch, and temporal information. This approach avoids error accumulation inherent in multi-stage methods and enhances accuracy through mutual reinforcement. Our framework employs tokenized representations specifically targeted for this task, through incorporating note value information. Furthermore, we introduce a pseudo-labeling technique to address a scarcity problem of annotated note value data. This technique produces approximate note value labels from existing datasets for the traditional note-level transcription. Experimental results demonstrate the superior performance of the proposed model in note-level transcription tasks when compared to existing state-of-the-art approaches. We also introduce new evaluation metrics that assess both temporal and note value aspects to demonstrate the robustness of the model. Moreover, qualitative assessments via visualized musical scores confirmed the effectiveness of our model in capturing the note values.
Problem

Research questions and friction points this paper is trying to address.

Transcribes audio to time-aligned musical scores
Extracts note values for accurate score generation
Uses end-to-end framework to enhance transcription accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

End-to-end framework integration
Tokenized note value representations
Pseudo-labeling for data scarcity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Leekyung Kim
Department of Industrial Engineering and Institute for Industrial Systems Innovation, Seoul National University, Seoul, Republic of Korea
S
Sungwook Jeon
Department of Industrial Engineering and Institute for Industrial Systems Innovation, Seoul National University, Seoul, Republic of Korea
W
Wan Heo
Department of Industrial Engineering and Institute for Industrial Systems Innovation, Seoul National University, Seoul, Republic of Korea
J
Jonghun Park
Department of Industrial Engineering and Institute for Industrial Systems Innovation, Seoul National University, Seoul, Republic of Korea