SepRQ : Self-Supervised Speech Mixture Representation Learning via Mask-Free, Multi-Scale Source Separation

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of mainstream self-supervised speech models in processing multi-speaker mixed audio by proposing SepRQ, an open-source framework. Departing from the conventional masked prediction paradigm, this method introduces a frozen random projection codebook and a multi-resolution pseudo-source separation mechanism, driving representation learning through a mask-agnostic, multi-scale pseudo-source separation objective. Experimental results demonstrate that SepRQ achieves state-of-the-art performance on both speaker diarization and speech separation tasks within the SUPERB benchmark. Notably, it comprehensively outperforms WavLM while utilizing fewer parameters, establishing a more efficient and effective approach for multi-speaker scenarios.
📝 Abstract
Self-supervised learning (SSL) is standard for speech representation learning, but mainstream models are designed around single-speaker audio, limiting their usefulness in multi-speakers scenarios. We present SepRQ, an open-source SSL framework that replaces masked prediction with a pseudo-source-separation objective over frozen random-projection codebooks. By adopting a novel mask-free, multiresolution approach, SepRQ achieves state-of-the-art performance in Speaker Diarization and Speech Separation on the SUPERB benchmark, surpassing WavLM and other cocktail-party derived SSLs at both Base and Large scales, while requiring only 85.68M inference parameters. SepRQ also demonstrates strong performance across target-speaker tasks requiring enrollment (such as Target-Speaker Automatic Speech Recognition), and on the challenging multi-domain DIHARD 3 diarization dataset. Notably, we report strong separation capabilities on three-speaker mixtures (WSJ0-3Mix), where current SSL literature struggles. While cocktail-party SSLs remain scarce and closed-source, limited to C-HuBERT and the enrollment-based SA-WavLM, we open-source SepRQ to the community.
Problem

Research questions and friction points this paper is trying to address.

Self-supervised learning
Speech mixture
Source separation
Speaker diarization
Multi-speaker scenarios
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-supervised learning
Source separation
Mask-free
Multi-scale representation
Speaker diarization
🔎 Similar Papers
No similar papers found.