Score
Designs, builds, and evaluates algorithms and processing pipelines that take mixed audio as input and produce per-source or enhanced audio outputs—this includes blind and supervised (including neural) source separation, speech enhancement, and audio restoration methods. Also develops preprocessing and separation-based classification workflows and evaluation metrics to handle overlapping voices or vocalizations, speaker-specific separation, and to measure or improve downstream detection/classification performance.
This work addresses the challenge of zero-shot extension of pretrained single-step audio source separation models to multi-step inference. We propose a gradient-free iterative framework that, at each step, linearly interpolates the mixture with the previous step’s output and adaptively selects the interpolation weight via optimization to maximize a separation metric. We provide the first theoretical proof that such single-step models can be elevated to multi-step systems through this procedure, establishing an equivalence to denoising diffusion bridges. Under model smoothness constraints, we derive a robustness error bound for the separation metric. Empirically, our method consistently outperforms single-step baselines on speech enhancement and music separation tasks; the performance gain matches that achieved by scaling model capacity, increasing training data, or performing explicit multi-step joint training. Crucially, improvements generalize across all evaluation metrics—even those not explicitly optimized during inference.
Current sound separation methods rely heavily on synthetically mixed training data, leading to limited generalization in real-world, naturally mixed acoustic scenarios. To address this, we propose ClearSep—a novel framework featuring a remix-based consistency dual-evaluation metric that drives joint optimization of separation and self-supervised distillation. ClearSep introduces an iterative data engine integrating self-supervised knowledge distillation, dynamic pseudo-label generation, and time-frequency masking-based separation. By enforcing remix consistency and adaptive thresholding, the framework enables customized training for individual source tracks. Evaluated across multiple benchmarks, ClearSep achieves state-of-the-art performance, significantly improving separation quality, robustness, and cross-domain generalization on naturally mixed audio—thereby overcoming the generalization bottleneck imposed by artificial mixing.
This work addresses the lack of a unified and reproducible experimental framework in music source separation research, which has hindered systematic comparison and rapid iteration. To this end, the authors propose MSST, an open-source framework featuring a YAML-driven architecture that integrates, for the first time, practical techniques such as sliding-window inference, test-time augmentation, model ensembling, and LoRA-based fine-tuning within a single pipeline. MSST supports diverse models, data augmentation strategies, loss functions, and evaluation metrics. Through comprehensive ablation studies, the authors demonstrate the effectiveness of these integrated components, showing consistent improvements in separation performance while substantially lowering the barrier to reproduction and development.
This work addresses the challenge of effectively separating impulsive acoustic events (e.g., knocks, alarms) from stationary background noise (e.g., HVAC hum, traffic rumble) in real-world soundscapes. We propose IS³, the first neural architecture specifically designed for impulse–stationary sound separation. IS³ integrates a lightweight deep neural network with a learnable deep filtering mechanism to achieve end-to-end, data-driven component disentanglement. To enhance generalizability across diverse acoustic sources, we introduce an efficient synthetic data generation pipeline that supports multi-source mixing and realistic spectral-temporal characteristics. Quantitative evaluation demonstrates that IS³ significantly outperforms conventional harmonic–percussive separation (HPS) and wavelet-based filtering methods across standard metrics (e.g., SI-SNR, SDR). These results validate the efficacy of learned separation paradigms in complex, non-stationary acoustic environments. The framework provides a robust foundation for downstream audio applications, including speech enhancement, adaptive noise suppression, and acoustic event detection.
Speech recognition and speaker change detection in smart glasses degrade significantly in noisy environments. Method: This paper proposes a multi-microphone directional speech enhancement method tailored for wearable devices, jointly modeling neural beamforming and multichannel source separation within an end-to-end optimized framework integrating separation and automatic speech recognition (ASR). The approach employs a lightweight Conv-TasNet variant and a differentiable beamformer to enable efficient directional source separation on resource-constrained edge devices. Contribution/Results: Directional separation alone reduces the word error rate (WER) of wearer’s speech by 32%. Joint training further enhances robustness, achieving state-of-the-art ASR performance for smart glasses under real-world noise conditions. Crucially, this work presents the first empirical validation of end-to-end joint optimization of source separation and ASR for wearable audio applications.
This study presents the first exploration into the detectability of AI-generated tracks within music co-created by humans and artificial intelligence. Addressing the challenge that general-purpose source separation methods often fail to reliably recover AI-specific artifacts, the authors propose a parallel detection architecture that operates without requiring full track separation. The approach leverages a neural audio codec to simulate the mixing process and combines short-time audio block analysis, relative energy estimation, and a binary classifier to directly identify AI-generated components within the mixed audio signal. Experiments on the MUSDB18-HQ dataset demonstrate promising track-level detection performance, confirming the feasibility of effectively discerning AI-generated content even in complex musical mixtures.
This work addresses the challenging task of recovering original, unprocessed stems from fully mixed and mastered music recordings. The proposed approach employs a two-stage pipeline: first, it aggregates outputs from multiple pre-trained source separation models to obtain initial stem estimates; second, it applies a dedicated BSRNN-based restoration model to each estimated stem to reverse complex audio degradations such as equalization, dynamic range compression, and reverberation. By uniquely combining ensemble-based separation with targeted restoration, the method effectively mitigates the compounded distortions present in real-world audio. Evaluated on the official MSR benchmark, the proposed system achieves the second-highest overall score, demonstrating substantial improvement over existing baselines.
This work addresses the challenge of accurately recovering individual watermarks from separated audio tracks in multi-track mixing and separation scenarios, where conventional watermarking methods often fail. To this end, we propose the first “separation-first” end-to-end joint training framework that simultaneously optimizes an audio separator and a multi-stream watermarking system. By integrating shared-structure multi-key watermark embedding, off-the-shelf separation model adaptation, and a joint training strategy, our approach enables watermark embedding to be robust to separation-induced distortions while encouraging the separator to preserve watermark-critical features. Experimental results demonstrate significant improvements in post-separation watermark bit recovery rates on both speech-plus-music and vocal-plus-accompaniment mixtures, all while maintaining high perceptual audio quality.
This work proposes a knowledge-driven, self-supervised approach to audio segmentation and source separation that circumvents the reliance on large-scale manually annotated data. By integrating external prior knowledge—such as musical scores—into the audio processing pipeline for the first time, the method leverages hidden Markov models to achieve effective segmentation and separation of music and film audio without requiring labeled training data. Evaluated on synthetic datasets, the approach demonstrates strong performance, and in real-world film soundtrack tests, it significantly outperforms purely data-driven methods when incorporating sound-class priors. This represents a notable advance toward annotation-free audio analysis through principled integration of domain knowledge.