Uncovering Ordinal-Matching Bias in Audio-Visual LLMs

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the frequent misalignment between speech and speakers by audio-visual large models in multi-speaker scenarios, revealing that this phenomenon stems from an "ordinal matching bias," whereby models rely on positional order rather than genuine audio-visual cues for association. To mitigate this issue, we propose Ordinal Decoupling Fine-Tuning (OD-FT), a strategy requiring only minimal synthetic data. This approach integrates diagnostic dataset construction with spatial-temporal randomized data augmentation to effectively rectify model biases. The proposed method significantly suppresses ordinal matching bias while enhancing generalization capabilities, achieving average performance improvements of 8.27% and 2.57% on Qwen2.5-Omni and video-SALMONN2+, respectively.
📝 Abstract
This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic dataset in which multiple visible speakers each utter a single word. Analysis on this corpus reveals a consistent error pattern across three recent open-source AVLLMs: models attribute utterances by simply matching the order of spoken sentences with the left-to-right, top-to-bottom arrangement of visible faces, rather than relying on audio-visual cues such as lip synchronization. We term this behavior \emph{ordinal-matching bias}. We further show that this bias can be substantially mitigated through a simple remedy, Ordinal-Decoupled Fine-Tuning (OD-FT), in which models are fine-tuned on synthetic videos where spatial positions of speakers and speaking order are independently randomized. Despite using only 400 synthetic training videos, OD-FT not only suppresses ordinal-matching bias but also improves audio-visual understanding on real-world videos, yielding average gains of 8.27\% for Qwen2.5-Omni and 2.57\% for video-SALMONN2+ across three audio-visual benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Audio-Visual Large Language Models
Ordinal-Matching Bias
Speaker Association
Multi-Speaker Scenes
Audio-Visual Understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio-Visual Large Language Models
Ordinal-Matching Bias
Ordinal-Decoupled Fine-Tuning
Synthetic Diagnostic Dataset
Speaker Identification
🔎 Similar Papers
No similar papers found.
J
Jihoo Jung
Korea Advanced Institute of Science and Technology, South Korea
Youngjoon Jang
Youngjoon Jang
KAIST
Computer VisionMachine Learning
H
Hyebin Cho
Korea Advanced Institute of Science and Technology, South Korea
S
Suho Yoo
Korea Advanced Institute of Science and Technology, South Korea
Joon Son Chung
Joon Son Chung
KAIST
Machine learningspeech processingcomputer vision