RadarVox: Radar-Audio Multimodal Cocktail-Party Speech Separation with Speaker-Aware Cross-Modal Matching

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the permutation ambiguity and speaker identity assignment challenges in blind audio source separation for cocktail party scenarios by proposing a radar-audio multimodal fusion solution. We construct a multimodal benchmark dataset and employ FMCW radar to capture laryngeal micro-movements, thereby obtaining spatial motion priors. A gated fusion DPRNN network is designed to achieve radar-assisted sound source separation, complemented by a cross-modal speaker-aware matcher that performs identity association. Experimental results demonstrate that the proposed method achieves an ordered SI-SDR improvement exceeding 13 dB and a speaker assignment accuracy above 80%, significantly outperforming audio-only baselines.
📝 Abstract
In embodied voice interaction, cocktail-party speech perception requires both speech separation and speaker attribution across machine-generated and human speech sources. However, conventional audio-only blind source separation remains permutation ambiguous, making the correspondence between separated streams and physical speakers unclear. This paper presents RadarVox, a radar-audio multimodal benchmark for identity-aware cocktail-party speech separation. RadarVox provides acoustic mixtures and source-level radar displacement signals from loudspeaker-emitted and human speech, enabling source-aware speaker assignment. FMCW radar captures laryngeal mechanical motion, providing speaker-specific spatial-motion cues unavailable to a single-channel microphone. We inject radar-derived priors into a DPRNN separator via gated fusion and learn a speaker-aware cross-modal matcher to associate unordered speech streams with radar-observed speakers. Experiments on multi-speaker mixtures show that radar cues provide complementary benefits, achieving scale-invariant signal-to-distortion ratio (SI-SDR) values of 9.75 dB and 7.13 dB in two- and three-speaker scenarios, respectively. More importantly, the proposed method improves ordered SI-SDR by more than 13 dB over audio-only methods while achieving over 80\% speaker assignment accuracy.
Problem

Research questions and friction points this paper is trying to address.

cocktail-party speech separation
permutation ambiguity
speaker attribution
radar-audio multimodal
embodied voice interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Radar-Audio Multimodal
Cocktail-Party Speech Separation
Cross-Modal Matching
Gated Fusion
DPRNN
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yanlin Xu
New Laboratory of Pattern Recognition (NLPR), State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences (CASIA), Beijing 100190, China; and School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS), Beijing 100049, China
Y
Yiwei Ru
Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China; and Beijing University of Posts and Telecommunications, Beijing 100876, China
M
Mupei Li
New Laboratory of Pattern Recognition (NLPR), State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences (CASIA), Beijing 100190, China; and School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS), Beijing 100049, China
Y
Yongji Liu
Beijing University of Posts and Telecommunications, Beijing 100876, China
J
Jie Wang
School of Information Science and Technology, Dalian Maritime University, Dalian 116026, China
Zhenan Sun
Zhenan Sun
Institute of Automation, Chinese Academy of Sciences
BiometricsPattern RecognitionComputer Vision