🤖 AI Summary
This study addresses the frequent misalignment between speech and speakers by audio-visual large models in multi-speaker scenarios, revealing that this phenomenon stems from an "ordinal matching bias," whereby models rely on positional order rather than genuine audio-visual cues for association. To mitigate this issue, we propose Ordinal Decoupling Fine-Tuning (OD-FT), a strategy requiring only minimal synthetic data. This approach integrates diagnostic dataset construction with spatial-temporal randomized data augmentation to effectively rectify model biases. The proposed method significantly suppresses ordinal matching bias while enhancing generalization capabilities, achieving average performance improvements of 8.27% and 2.57% on Qwen2.5-Omni and video-SALMONN2+, respectively.
📝 Abstract
This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic dataset in which multiple visible speakers each utter a single word. Analysis on this corpus reveals a consistent error pattern across three recent open-source AVLLMs: models attribute utterances by simply matching the order of spoken sentences with the left-to-right, top-to-bottom arrangement of visible faces, rather than relying on audio-visual cues such as lip synchronization. We term this behavior \emph{ordinal-matching bias}. We further show that this bias can be substantially mitigated through a simple remedy, Ordinal-Decoupled Fine-Tuning (OD-FT), in which models are fine-tuned on synthetic videos where spatial positions of speakers and speaking order are independently randomized. Despite using only 400 synthetic training videos, OD-FT not only suppresses ordinal-matching bias but also improves audio-visual understanding on real-world videos, yielding average gains of 8.27\% for Qwen2.5-Omni and 2.57\% for video-SALMONN2+ across three audio-visual benchmarks.