audio source separation

Designs, builds, and evaluates algorithms and processing pipelines that take mixed audio as input and produce per-source or enhanced audio outputs—this includes blind and supervised (including neural) source separation, speech enhancement, and audio restoration methods. Also develops preprocessing and separation-based classification workflows and evaluation metrics to handle overlapping voices or vocalizations, speaker-specific separation, and to measure or improve downstream detection/classification performance.

audiosourceseparation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Training-Free Multi-Step Audio Source Separation

May 26, 2025
YZ
Yongyi Zang
🏛️ University of Illinois Urbana-Champaign | The Chinese University of Hong Kong

This work addresses the challenge of zero-shot extension of pretrained single-step audio source separation models to multi-step inference. We propose a gradient-free iterative framework that, at each step, linearly interpolates the mixture with the previous step’s output and adaptively selects the interpolation weight via optimization to maximize a separation metric. We provide the first theoretical proof that such single-step models can be elevated to multi-step systems through this procedure, establishing an equivalence to denoising diffusion bridges. Under model smoothness constraints, we derive a robustness error bound for the separation metric. Empirically, our method consistently outperforms single-step baselines on speech enhancement and music separation tasks; the performance gain matches that achieved by scaling model capacity, increasing training data, or performing explicit multi-step joint training. Crucially, improvements generalize across all evaluation metrics—even those not explicitly optimized during inference.

Enhancing audio separation without additional trainingIterative blending for multi-step source separationTheoretical link to denoising diffusion bridge models

Unleashing the Power of Natural Audio Featuring Multiple Sound Sources

Apr 24, 2025
XC
Xize Cheng
🏛️ Zhejiang University | Independent Researcher

Current sound separation methods rely heavily on synthetically mixed training data, leading to limited generalization in real-world, naturally mixed acoustic scenarios. To address this, we propose ClearSep—a novel framework featuring a remix-based consistency dual-evaluation metric that drives joint optimization of separation and self-supervised distillation. ClearSep introduces an iterative data engine integrating self-supervised knowledge distillation, dynamic pseudo-label generation, and time-frequency masking-based separation. By enforcing remix consistency and adaptive thresholding, the framework enables customized training for individual source tracks. Evaluated across multiple benchmarks, ClearSep achieves state-of-the-art performance, significantly improving separation quality, robustness, and cross-domain generalization on naturally mixed audio—thereby overcoming the generalization bottleneck imposed by artificial mixing.

Generalization to real-world natural audioQuantitative evaluation of separation qualityUniversal sound separation from mixed audio

This work addresses the lack of a unified and reproducible experimental framework in music source separation research, which has hindered systematic comparison and rapid iteration. To this end, the authors propose MSST, an open-source framework featuring a YAML-driven architecture that integrates, for the first time, practical techniques such as sliding-window inference, test-time augmentation, model ensembling, and LoRA-based fine-tuning within a single pipeline. MSST supports diverse models, data augmentation strategies, loss functions, and evaluation metrics. Through comprehensive ablation studies, the authors demonstrate the effectiveness of these integrated components, showing consistent improvements in separation performance while substantially lowering the barrier to reproduction and development.

Demixing ModelsEvaluation MetricsMusic Source Separation

IS${}^3$ : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering

Sep 01, 2025
CB
Clémentine Berger
🏛️ Télécom Paris | Institut Polytechnique de Paris

This work addresses the challenge of effectively separating impulsive acoustic events (e.g., knocks, alarms) from stationary background noise (e.g., HVAC hum, traffic rumble) in real-world soundscapes. We propose IS³, the first neural architecture specifically designed for impulse–stationary sound separation. IS³ integrates a lightweight deep neural network with a learnable deep filtering mechanism to achieve end-to-end, data-driven component disentanglement. To enhance generalizability across diverse acoustic sources, we introduce an efficient synthetic data generation pipeline that supports multi-source mixing and realistic spectral-temporal characteristics. Quantitative evaluation demonstrates that IS³ significantly outperforms conventional harmonic–percussive separation (HPS) and wavelet-based filtering methods across standard metrics (e.g., SI-SNR, SDR). These results validate the efficacy of learned separation paradigms in complex, non-stationary acoustic environments. The framework provides a robust foundation for downstream audio applications, including speech enhancement, adaptive noise suppression, and acoustic event detection.

Enable differentiated audio processing for applicationsProvide pre-processing for noise suppression and classificationSeparate impulsive sounds from stationary background

Directional Source Separation for Robust Speech Recognition on Smart Glasses

Sep 20, 2023
TF
Tiantian Feng
🏛️ University of Southern California | Meta Platforms Inc.

Speech recognition and speaker change detection in smart glasses degrade significantly in noisy environments. Method: This paper proposes a multi-microphone directional speech enhancement method tailored for wearable devices, jointly modeling neural beamforming and multichannel source separation within an end-to-end optimized framework integrating separation and automatic speech recognition (ASR). The approach employs a lightweight Conv-TasNet variant and a differentiable beamformer to enable efficient directional source separation on resource-constrained edge devices. Contribution/Results: Directional separation alone reduces the word error rate (WER) of wearer’s speech by 32%. Joint training further enhances robustness, achieving state-of-the-art ASR performance for smart glasses under real-world noise conditions. Crucially, this work presents the first empirical validation of end-to-end joint optimization of source separation and ASR for wearable audio applications.

Enhance directional source separation using multi-microphone arraysImprove speech recognition in noisy environmentsOptimize joint training of separation and ASR models

Latest Papers

What's happening recently
View more

This study presents the first exploration into the detectability of AI-generated tracks within music co-created by humans and artificial intelligence. Addressing the challenge that general-purpose source separation methods often fail to reliably recover AI-specific artifacts, the authors propose a parallel detection architecture that operates without requiring full track separation. The approach leverages a neural audio codec to simulate the mixing process and combines short-time audio block analysis, relative energy estimation, and a binary classifier to directly identify AI-generated components within the mixed audio signal. Experiments on the MUSDB18-HQ dataset demonstrate promising track-level detection performance, confirming the feasibility of effectively discerning AI-generated content even in complex musical mixtures.

AI-generated stemshybrid human-AI musicmusic detection

This work addresses the challenging task of recovering original, unprocessed stems from fully mixed and mastered music recordings. The proposed approach employs a two-stage pipeline: first, it aggregates outputs from multiple pre-trained source separation models to obtain initial stem estimates; second, it applies a dedicated BSRNN-based restoration model to each estimated stem to reverse complex audio degradations such as equalization, dynamic range compression, and reverberation. By uniquely combining ensemble-based separation with targeted restoration, the method effectively mitigates the compounded distortions present in real-world audio. Evaluated on the official MSR benchmark, the proposed system achieves the second-highest overall score, demonstrating substantial improvement over existing baselines.

Audio RestorationMastered MusicMusic Source Restoration

This work addresses the challenge of accurately recovering individual watermarks from separated audio tracks in multi-track mixing and separation scenarios, where conventional watermarking methods often fail. To this end, we propose the first “separation-first” end-to-end joint training framework that simultaneously optimizes an audio separator and a multi-stream watermarking system. By integrating shared-structure multi-key watermark embedding, off-the-shelf separation model adaptation, and a joint training strategy, our approach enables watermark embedding to be robust to separation-induced distortions while encouraging the separator to preserve watermark-critical features. Experimental results demonstrate significant improvements in post-separation watermark bit recovery rates on both speech-plus-music and vocal-plus-accompaniment mixtures, all while maintaining high perceptual audio quality.

audio watermarkingmulti-streamseparation artifacts

This work proposes a knowledge-driven, self-supervised approach to audio segmentation and source separation that circumvents the reliance on large-scale manually annotated data. By integrating external prior knowledge—such as musical scores—into the audio processing pipeline for the first time, the method leverages hidden Markov models to achieve effective segmentation and separation of music and film audio without requiring labeled training data. Evaluated on synthetic datasets, the approach demonstrates strong performance, and in real-world film soundtrack tests, it significantly outperforms purely data-driven methods when incorporating sound-class priors. This represents a notable advance toward annotation-free audio analysis through principled integration of domain knowledge.

audio source separationcinematic audioknowledge-driven

Hot Scholars

YH

Yuan-Hao Wei

Hong Kong Polytechnic University
Machine LearningVAEICA
HL

Haizhou Li

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China; NUS, Singapore
Automatic Speech RecognitionSpeaker RecognitionLanguage RecognitionVoice Conversion
HY

Han Yin

Tongyi Speech Lab, Alibaba Group
Audio UnderstandingMultimodal LLM
TV

Tuomas Virtanen

Tampere University
machine listeningaudio signal processingaudio