🤖 AI Summary
This work addresses key challenges in single-channel speech separation—namely, source permutation ambiguity, unstable sampling in generative models, and difficulties in aligning long audio segments across chunks—by proposing an ordered two-speaker separation method based on conditional flow matching. The approach fixes the source ordering during training by freezing the speaker encoder and introduces a biometric-guided Best-of-N sampling mechanism together with a chunk alignment strategy during inference to ensure stable and consistently ordered outputs. Built upon a Transformer U-Net architecture, the method achieves competitive performance on Libri2Mix in terms of SI-SDR, PESQ, and ESTOI metrics, while substantially reducing error rates in downstream tasks such as automatic speech recognition (cpWER) and speaker verification (EER).
📝 Abstract
Single-channel speech separation remains challenging for real-world deployment due to source permutation ambiguity, sampling variability of generative models, and the difficulty of processing long recordings with chunk-wise inference. We address these issues with a conditional flow-matching-based method that produces an ordered two-source output conditioned on the mixture. A frozen speaker encoder defines the source order during training and is reused at inference for biometric best-of-$N$ candidate selection and chunk-level channel alignment. We evaluate separation quality on Libri2Mix benchmark using SI-SDR, PESQ, and ESTOI, and measure downstream impact using cpWER for automatic speech recognition and EER for speaker verification. The results show that the proposed Transformer U-Net variant is competitive with strong baselines in objective separation metrics and achieves the lowest downstream automatic speech recognition and speaker verification error rates in all evaluated settings.