On a Separate Note: Robust Score-Informed Note Separation with a Two-Stream TFC-TDF U-Net and Adaptive Set Ownership

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing deep learning audio separation systems, which support only instrument-level track extraction and fail to achieve score-based, note-level waveform reconstruction. We propose NoteSep, the first score-guided, note-level separation framework, employing a dual-stream TFC-TDF U-Net architecture with bidirectional cross-attention. The method introduces selective harmonic gating to suppress low-frequency interference while preserving percussive transients, combined with an adaptive set ownership algorithm to resolve note attribution in polyphonic scenarios. Evaluated on the SCNS-Eval dataset, NoteSep achieves a median SI-SDR of 7.39 dB, substantially outperforming the strongest baseline at 2.49 dB. This work represents a significant breakthrough in precisely extracting individual notes from polyphonic recordings.
📝 Abstract
Score-informed note separation seeks to extract the performed waveform of all individual notes, often from a polyphonic recording. Existing deep learning systems generally only target instrument-level stems. We present, to our knowledge, the first deep learning approach to score-informed note separation, NoteSep. NoteSep extracts the queried notes by applying an extraction stage model, NoteGrab, once per note. Conditioned on pitch, onset, and offset, NoteGrab separates harmonic and percussive components in two U-Nets linked by bidirectional cross-attention; selective harmonic gating suppresses lower-octave interference while preserving percussive attacks. Finally, a joint separation stage applies Adaptive Set Ownership (ASO) to compare concurrent NoteGrab estimates and reallocate mixture energy. We curate SCNS-Train (25,729 mixtures and 743,920 targets) for training and SCNS-Eval (16 instruments, disjoint scores and libraries) for evaluation. On SCNS-Eval, NoteSep reaches a median SI-SDR of 7.39~dB, compared with 2.49~dB for our strongest baseline. See the demo page at https://benschou.com/notesep.
Problem

Research questions and friction points this paper is trying to address.

score-informed note separation
polyphonic music
source separation
deep learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Score-Informed Note Separation
Two-Stream TFC-TDF U-Net
Bidirectional Cross-Attention
Adaptive Set Ownership
Selective Harmonic Gating
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.