🤖 AI Summary
This paper addresses the challenging problem of multi-source separation in a cappella music, where the number of singers varies dynamically over time. To tackle this, we propose SepACap—a novel end-to-end separation model. Methodologically: (1) we design a power-set-based data augmentation strategy to comprehensively cover diverse singer configurations; (2) we introduce periodic activation functions to improve robustness during silent segments; (3) we formulate a composite loss function to enhance generalization across variable numbers of singers; and (4) we adopt SepReformer as the backbone architecture, integrated with spectrogram-domain transformations. Evaluated on the JaCappella dataset, SepACap significantly outperforms conventional spectrogram-based baselines, achieving state-of-the-art performance on both full-combination and subset separation tasks. To our knowledge, it is the first method to enable high-accuracy, generalizable separation of dynamically varying a cappella ensembles.
📝 Abstract
In this work, we study the task of multi-singer separation in a cappella music, where the number of active singers varies across mixtures. To address this, we use a power set-based data augmentation strategy that expands limited multi-singer datasets into exponentially more training samples. To separate singers, we introduce SepACap, an adaptation of SepReformer, a state-of-the-art speaker separation model architecture. We adapt the model with periodic activations and a composite loss function that remains effective when stems are silent, enabling robust detection and separation. Experiments on the JaCappella dataset demonstrate that our approach achieves state-of-the-art performance in both full-ensemble and subset singer separation scenarios, outperforming spectrogram-based baselines while generalizing to realistic mixtures with varying numbers of singers.