ROAM-ASD: Robust Open-World Active Speaker Detection with Flexible Multimodal Fusion

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出ROAM-ASD框架,通过联合建模音频、全脸和嘴部细粒度表示及多模态融合方法解决复杂条件下主动说话人检测准确性下降的问题。
📝 Abstract
Active speaker detection (ASD) requires reliable association between visible faces and acoustic speech, yet existing systems often degrade under challenging domains or incomplete observations. We introduce ROAM-ASD, a robust audiovisual framework that jointly models audio, full-face, and fine-grained mouth representations. A unified joint self-attention mechanism processes all input streams together with modality-agnostic query tokens, enabling direct interaction among available modality inputs. Modality dropout further improves robustness when input streams are unavailable. ROAM-ASD achieves state-of-the-art performance across five ASD benchmarks: 98.8% mAP on WASD, 87.9% on UniTalk, 96.5% on AVA, 99.3% on ASW, and 98.2% on Talkies, improving over previous best systems by 5.1, 4.7, 0.9, 1.0, and 2.1 mAP points, respectively. ROAM-ASD also substantially improves zero-shot cross-dataset generalization and remains robust to missing observations.
Problem

Research questions and friction points this paper is trying to address.

Active Speaker Detection
Challenging Domains
Incomplete Observations
Audiovisual Framework
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Joint Self-Attention
Modality Dropout
Multimodal Fusion
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pu Wang
KU Leuven, Department of Electrical Engineering, Leuven, Belgium
Yujun Wang
Yujun Wang
AIP Publishing
Theoretical AtomicMolecularand Optical Physics
H
Hugo Van hamme
KU Leuven, Department of Electrical Engineering, Leuven, Belgium