🤖 AI Summary
This study addresses the permutation ambiguity and speaker identity assignment challenges in blind audio source separation for cocktail party scenarios by proposing a radar-audio multimodal fusion solution. We construct a multimodal benchmark dataset and employ FMCW radar to capture laryngeal micro-movements, thereby obtaining spatial motion priors. A gated fusion DPRNN network is designed to achieve radar-assisted sound source separation, complemented by a cross-modal speaker-aware matcher that performs identity association. Experimental results demonstrate that the proposed method achieves an ordered SI-SDR improvement exceeding 13 dB and a speaker assignment accuracy above 80%, significantly outperforming audio-only baselines.
📝 Abstract
In embodied voice interaction, cocktail-party speech perception requires both speech separation and speaker attribution across machine-generated and human speech sources. However, conventional audio-only blind source separation remains permutation ambiguous, making the correspondence between separated streams and physical speakers unclear. This paper presents RadarVox, a radar-audio multimodal benchmark for identity-aware cocktail-party speech separation. RadarVox provides acoustic mixtures and source-level radar displacement signals from loudspeaker-emitted and human speech, enabling source-aware speaker assignment. FMCW radar captures laryngeal mechanical motion, providing speaker-specific spatial-motion cues unavailable to a single-channel microphone. We inject radar-derived priors into a DPRNN separator via gated fusion and learn a speaker-aware cross-modal matcher to associate unordered speech streams with radar-observed speakers. Experiments on multi-speaker mixtures show that radar cues provide complementary benefits, achieving scale-invariant signal-to-distortion ratio (SI-SDR) values of 9.75 dB and 7.13 dB in two- and three-speaker scenarios, respectively. More importantly, the proposed method improves ordered SI-SDR by more than 13 dB over audio-only methods while achieving over 80\% speaker assignment accuracy.