Source Separation for A Cappella Music

📅 2025-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the challenging problem of multi-source separation in a cappella music, where the number of singers varies dynamically over time. To tackle this, we propose SepACap—a novel end-to-end separation model. Methodologically: (1) we design a power-set-based data augmentation strategy to comprehensively cover diverse singer configurations; (2) we introduce periodic activation functions to improve robustness during silent segments; (3) we formulate a composite loss function to enhance generalization across variable numbers of singers; and (4) we adopt SepReformer as the backbone architecture, integrated with spectrogram-domain transformations. Evaluated on the JaCappella dataset, SepACap significantly outperforms conventional spectrogram-based baselines, achieving state-of-the-art performance on both full-combination and subset separation tasks. To our knowledge, it is the first method to enable high-accuracy, generalizable separation of dynamically varying a cappella ensembles.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Multi-modal VisionCognitive Modeling & Cognitive Systems: Affective Computing

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Accountability, Transparency, and Ethics for personalization
📝 Abstract
In this work, we study the task of multi-singer separation in a cappella music, where the number of active singers varies across mixtures. To address this, we use a power set-based data augmentation strategy that expands limited multi-singer datasets into exponentially more training samples. To separate singers, we introduce SepACap, an adaptation of SepReformer, a state-of-the-art speaker separation model architecture. We adapt the model with periodic activations and a composite loss function that remains effective when stems are silent, enabling robust detection and separation. Experiments on the JaCappella dataset demonstrate that our approach achieves state-of-the-art performance in both full-ensemble and subset singer separation scenarios, outperforming spectrogram-based baselines while generalizing to realistic mixtures with varying numbers of singers.
Problem

Research questions and friction points this paper is trying to address.

Separating multiple singers in a cappella music
Handling variable numbers of active singers
Robust detection and separation with silent stems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Power set data augmentation expands limited training samples
SepACap adapts SepReformer with periodic activation functions
Composite loss function handles silent stems robustly
🔎 Similar Papers
No similar papers found.