🤖 AI Summary
This study addresses the challenge of instrument source separation in orchestral recordings, where pitch, harmonics, and timbre are highly overlapping. We propose a score-guided deep neural network that leverages musical scores to construct frame-wise query vectors for matching shared audio representations, recovering individual instrument parts through joint time-frequency evidence allocation via Softmax. Additionally, a score-conditioning mechanism is introduced to preserve instrument identity even when notes are missing from the score. Experimental results demonstrate that the proposed method achieves state-of-the-art average signal-to-distortion ratio (SDR) on real-world recordings. Furthermore, it exhibits strong zero-shot generalization capabilities and significant robustness to incomplete or corrupted musical scores.
📝 Abstract
Orchestral separation recovers instrument sections from mixtures in which shared pitches, harmonics, and timbres obscure source identity. An aligned score provides instrument labels, note pitches, and activity times. A score-informed approach appends piano rolls to audio features before mask prediction. We introduce SCISSOR (Score-Conditioned Instrument Source Separation for Orchestral Recordings), which uses the score to form a frame-wise query for each source. Each query matches a shared audio representation, and a softmax over instrument and background slots jointly assigns overlapping time-frequency evidence. The queries retain instrument identity even when notes are missing from the score. After training on SynthSOD and a small set of URMP and PHENICX-Anechoic recordings, SCISSOR achieves the highest average SDR on held-out real recordings. With SynthSOD-only training, it leads on SynthSOD and zero-shot PHENICX-Anechoic, and improves on its audio-only control on zero-shot URMP. SCISSOR also degrades less under score corruption than the evaluated score-based baselines.