Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling

📅 2026-07-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses key challenges in single-channel speech separation—namely, source permutation ambiguity, unstable sampling in generative models, and difficulties in aligning long audio segments across chunks—by proposing an ordered two-speaker separation method based on conditional flow matching. The approach fixes the source ordering during training by freezing the speaker encoder and introduces a biometric-guided Best-of-N sampling mechanism together with a chunk alignment strategy during inference to ensure stable and consistently ordered outputs. Built upon a Transformer U-Net architecture, the method achieves competitive performance on Libri2Mix in terms of SI-SDR, PESQ, and ESTOI metrics, while substantially reducing error rates in downstream tasks such as automatic speech recognition (cpWER) and speaker verification (EER).
📝 Abstract
Single-channel speech separation remains challenging for real-world deployment due to source permutation ambiguity, sampling variability of generative models, and the difficulty of processing long recordings with chunk-wise inference. We address these issues with a conditional flow-matching-based method that produces an ordered two-source output conditioned on the mixture. A frozen speaker encoder defines the source order during training and is reused at inference for biometric best-of-$N$ candidate selection and chunk-level channel alignment. We evaluate separation quality on Libri2Mix benchmark using SI-SDR, PESQ, and ESTOI, and measure downstream impact using cpWER for automatic speech recognition and EER for speaker verification. The results show that the proposed Transformer U-Net variant is competitive with strong baselines in objective separation metrics and achieves the lowest downstream automatic speech recognition and speaker verification error rates in all evaluated settings.
Problem

Research questions and friction points this paper is trying to address.

speech separation
permutation ambiguity
generative sampling variability
long-form audio processing
single-channel
Innovation

Methods, ideas, or system contributions that make the work stand out.

flow matching
speech source separation
biometric sampling
best-of-N selection
channel alignment
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
A
Anastasia Zorkina
ITMO University, Speech Processing Group, Russia
A
Alexandr Anikin
ITMO University, Speech Processing Group, Russia
N
Nikita Khmelev
ITMO University, Speech Processing Group, Russia; Speech Technology Center Ltd., R&D department, Russia
A
Anastasiya Korenevskaya
ITMO University, Speech Processing Group, Russia
Sergey Novoselov
Sergey Novoselov
Speech Technology Center Ltd.
machine learningspeaker recognitionlanguage recognitionspeech signal processing
V
Vladimir Volokhov
ITMO University, Speech Processing Group, Russia; Speech Technology Center Ltd., R&D department, Russia
M
Maxim Korenevsky
Speech Technology Center Ltd., R&D department, Russia
Y
Yuriy Matveev
ITMO University, Speech Processing Group, Russia