🤖 AI Summary
Traditional music source separation (MSS) methods are constrained by predefined instrument taxonomies and lack the ability to separate arbitrary target sources specified via audio queries. To address this, we propose QSCNet—a query-driven separation framework based on a conditional U-Net architecture. QSCNet conditions separation on a raw audio snippet of the target source, integrating an embedded audio query mechanism with a Sparse Compression Network (SCN) to jointly model query-target relationships. This enables, for the first time, high-fidelity conditional separation with a U-Net backbone without reliance on fixed instrument vocabularies. Evaluated on MoisesDB, QSCNet achieves a 1.0 dB improvement in signal-to-noise ratio (SNR) over the state-of-the-art Banquet model, while using less than 50% of its parameters—demonstrating superior trade-offs between separation accuracy and computational efficiency.
📝 Abstract
In this paper we propose a conditioned UNet for Music Source Separation (MSS). MSS is generally performed by multi-output neural networks, typically UNets, with each output representing a particular stem from a predefined instrument vocabulary. In contrast, conditioned MSS networks accept an audio query related to a stem of interest alongside the signal from which that stem is to be extracted. Thus, a strict vocabulary is not required and this enables more realistic tasks in MSS. The potential of conditioned approaches for such tasks has been somewhat hidden due to a lack of suitable data, an issue recently addressed with the MoisesDb dataset. A recent method, Banquet, employs this dataset with promising results seen on larger vocabularies. Banquet uses Bandsplit RNN rather than a UNet and the authors state that UNets should not be suitable for conditioned MSS. We counter this argument and propose QSCNet, a novel conditioned UNet for MSS that integrates network conditioning elements in the Sparse Compressed Network for MSS. We find QSCNet to outperform Banquet by over 1dB SNR on a couple of MSS tasks, while using less than half the number of parameters.