Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

📅 2026-07-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of a unified and reproducible experimental framework in music source separation research, which has hindered systematic comparison and rapid iteration. To this end, the authors propose MSST, an open-source framework featuring a YAML-driven architecture that integrates, for the first time, practical techniques such as sliding-window inference, test-time augmentation, model ensembling, and LoRA-based fine-tuning within a single pipeline. MSST supports diverse models, data augmentation strategies, loss functions, and evaluation metrics. Through comprehensive ablation studies, the authors demonstrate the effectiveness of these integrated components, showing consistent improvements in separation performance while substantially lowering the barrier to reproduction and development.
📝 Abstract
Music Source Separation (MSS), the task of recovering individual sound components (stems) from a polyphonic mixture, is central to applications ranging from karaoke and remixing to audio restoration and content production. The separation quality depends on engineering decisions across the entire pipeline: model choice, training data preparation and augmentation, loss function and metrics choice, training configuration, validation, and post-processing. This paper presents MSST (Music-Source-Separation-Training) - a universal open-source framework for MSS tasks, which unifies training, validation, and inference for a broad range of modern demixing model families under a single, configuration-driven interface. The framework supports various model architectures, data preprocessing and augmentations, multiple loss functions and evaluation metrics, which helps with fast iterations and ablation studies. Additionally, the framework supports a range of practical techniques that improve separation quality, such as sliding-window inference with cross-fading, test-time augmentation, model ensembling, and fine-tuning via Low-Rank Adaptation (LORA). Our ablation studies demonstrate improvements of MSS using the above techniques. By consolidating these components into a reproducible, YAML-configurable framework, MSST lowers the barrier to systematic experimentation and enables rapid iteration from idea to verifiable result.
Problem

Research questions and friction points this paper is trying to address.

Music Source Separation
Demixing Models
Training Framework
Evaluation Metrics
Systematic Experimentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Music Source Separation
Unified Framework
Low-Rank Adaptation (LoRA)
Test-Time Augmentation
Sliding-Window Inference
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Roman Solovyev
MVSep.com
I
Ilya Kiselev
National Research University Higher School of Economics (HSE University), MIEM HSE
A
Alexander Stempkovskiy
AlphaChip LLC.
T
Tatiana Gabruseva
Independent Researcher