Toward Part-Aware Choral Transcription with singing voice assignment

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of voice overlap and the difficulty of separating independent SATB tracks in choral music transcription by proposing the first end-to-end neural framework that jointly models note transcription and voice assignment. Methodologically, it introduces a novel voice-aware transcription paradigm utilizing neural networks to predict note onset and offset frames. Furthermore, it incorporates vocal range priors and ordered continuity to construct a structured training objective, combined with a voice presence estimation strategy to achieve precise transcription. Experimental results demonstrate that the proposed method improves the voice-aware F1 score on the YouChorale dataset by 36.4% over baseline models, significantly enhancing the separation performance of individual tracks from polyphonic choral audio.
📝 Abstract
Single-instrument automatic music transcription (AMT) has advanced substantially, yet choral applications require soprano, alto, tenor, and bass (SATB) to be transcribed as separate parts. Recent note-level choral AMT instead produces a single merged note track, limiting rehearsal, education, and score reconstruction. To address this limitation, we introduce Part-aware Choral Transcription (PawCT), to our knowledge the first end-to-end neural framework that identifies active SATB parts from choral audio and transcribes each into a separate note-level track. PawCT combines part-specific onset, offset, and frame prediction with part-presence estimation, union-level supervision, and structured training targets using a range prior (RP) based on SATB pitch ranges and its ordered-continuity (OC) extension, which adds within-part melodic continuity and cross-part pitch ordering. On YouChorale, PawCT-RP-OC achieves a macro part-aware note F1 of 0.225 at a 50-ms onset tolerance, outperforming an adapted choral baseline (0.165) by 36.4% relative and a two-stage post-hoc assignment pipeline (0.175). Its part-agnostic variant, PagCT, achieves a 50-ms onset F1 of 0.382, compared with 0.237 for the previous state-of-the-art choral AMT model. Cross-dataset evaluations on CSD and Cantoria further assess performance under dataset shift. These results demonstrate the benefit of jointly modeling note transcription and vocal-part assignment. Code and demos are available at https://hanyu-meng.github.io/Paw_Choral_AMT_Demo/.
Problem

Research questions and friction points this paper is trying to address.

Choral Transcription
Automatic Music Transcription
Singing Voice Assignment
Part-aware
SATB
Innovation

Methods, ideas, or system contributions that make the work stand out.

Choral Transcription
End-to-end Framework
Singing Voice Assignment
Range Prior
Ordered-Continuity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hanyu Meng
The University of New South Wales, Sydney, Australia
Zhanhong He
Zhanhong He
PhD Student, University of Western Australia
Automatic Music TranscriptionAudio Processing
Z
Zixun Guo
Center for Digital Music (C4DM), Queen Mary University of London, London, United Kingdom
Yaolong Ju
Yaolong Ju
Great Bay University
Automatic Music TranscriptionMachine LearningMusic TheoryMusic Information Retrieval