A Conditioned UNet for Music Source Separation

📅 2025-12-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional music source separation (MSS) methods are constrained by predefined instrument taxonomies and lack the ability to separate arbitrary target sources specified via audio queries. To address this, we propose QSCNet—a query-driven separation framework based on a conditional U-Net architecture. QSCNet conditions separation on a raw audio snippet of the target source, integrating an embedded audio query mechanism with a Sparse Compression Network (SCN) to jointly model query-target relationships. This enables, for the first time, high-fidelity conditional separation with a U-Net backbone without reliance on fixed instrument vocabularies. Evaluated on MoisesDB, QSCNet achieves a 1.0 dB improvement in signal-to-noise ratio (SNR) over the state-of-the-art Banquet model, while using less than 50% of its parameters—demonstrating superior trade-offs between separation accuracy and computational efficiency.

Technology Category

Machine Learning: Mixture of Experts (MoE)Cognitive Modeling & Cognitive Systems: Neural Spike CodingSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsGraph Algorithms and Modeling for the Web: Querying, indexing, and retrieval in Web-related graphs
📝 Abstract
In this paper we propose a conditioned UNet for Music Source Separation (MSS). MSS is generally performed by multi-output neural networks, typically UNets, with each output representing a particular stem from a predefined instrument vocabulary. In contrast, conditioned MSS networks accept an audio query related to a stem of interest alongside the signal from which that stem is to be extracted. Thus, a strict vocabulary is not required and this enables more realistic tasks in MSS. The potential of conditioned approaches for such tasks has been somewhat hidden due to a lack of suitable data, an issue recently addressed with the MoisesDb dataset. A recent method, Banquet, employs this dataset with promising results seen on larger vocabularies. Banquet uses Bandsplit RNN rather than a UNet and the authors state that UNets should not be suitable for conditioned MSS. We counter this argument and propose QSCNet, a novel conditioned UNet for MSS that integrates network conditioning elements in the Sparse Compressed Network for MSS. We find QSCNet to outperform Banquet by over 1dB SNR on a couple of MSS tasks, while using less than half the number of parameters.
Problem

Research questions and friction points this paper is trying to address.

Proposes a conditioned UNet for music source separation tasks
Addresses the limitation of predefined instrument vocabularies in separation
Demonstrates improved performance with fewer parameters than existing methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conditioned UNet for music source separation
Integrates conditioning in Sparse Compressed Network
Outperforms prior method with fewer parameters
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Ken O'Hanlon
Centre For Digital Music, Queen Mary University of London
B
Basil Woods
AudioStrip Ltd. London
L
Lin Wang
Centre For Digital Music, Queen Mary University of London
M
Mark Sandler
Centre For Digital Music, Queen Mary University of London