Beyond Text Following: Repairable Arbitration Reversals in Audio-Language Models

📅 2026-06-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the tendency of audio language models to blindly follow erroneous textual inputs even when confronted with clear and contradictory audio evidence. Through a counterfactual experimental setup that fixes the audio while removing conflicting text, the study reveals that although these models encode relevant acoustic information, it is systematically overridden by textual cues during cross-modal arbitration. To mitigate this bias, the authors propose GACL (Gated Audio Counterfactual Logic), a training-free decoding rule that leverages counterfactual reasoning to recalibrate modality trust. Evaluated under strict fidelity constraints—allowing no more than a 5-percentage-point drop in faithfulness—GACL improves normalized Area Under the Curve (nAUC) by 17.8 points over the strongest baseline. Remarkably, without any task-specific tuning, GACL generalizes to vision–language arbitration tasks, yielding gains of up to 40.5 percentage points.
📝 Abstract
Audio-language models (ALMs) often follow text that conflicts with audio, even when the audio evidence is clear. This raises a basic question: is the audio-supported answer unavailable, or is it represented but overridden by the conflicting text? We examine this question using a same-audio counterfactual that keeps the audio fixed, removes only the conflicting text, and measures the resulting shift in model preference. Across five ALMs and four conflict tasks, 64.1% of conflict samples show a sign flip: the same-audio branch prefers the audio-supported answer, whereas the joint branch prefers the text-supported answer. This pattern suggests that the relevant audio evidence is encoded but loses in arbitration. Activation patching further localizes the reversal to answer-position computation, and patching effects closely track output candidate-score differences (Spearman rho=0.93). Using this diagnostic, we propose Gated Audio Counterfactual Logit Correction (GACL), a training-free decoding rule that interpolates between joint and same-audio scores. Under a strict 5 pp faithfulness-drop budget, GACL improves nAUC by 17.8 points over the best contrastive baseline and transfers without retuning to vision-text arbitration (up to +40.5 pp).
Problem

Research questions and friction points this paper is trying to address.

audio-language models
multimodal arbitration
text-audio conflict
model faithfulness
counterfactual analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio-language models
counterfactual analysis
activation patching
arbitration reversal
training-free decoding
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yichen Gao
Northeastern University, China
Yiqun Zhang
Yiqun Zhang
Northeastern University, China; Shanghai Artificial Intelligence Laboratory
empathetic dialogueLLM-based agentMulti-agent
Z
Zijing Wang
Northeastern University, China
Yujia Li
Yujia Li
Research Scientist, Google DeepMind
Machine LearningComputer VisionNatural Language ProcessingOptimization
H
Heng Guo
Northeastern University, China
X
Xi Wu
Northeastern University, China
Xiaocui Yang
Xiaocui Yang
Lecturer, Northeastern University (China)
Multimodal Sentiment AnalysisData MiningMultimodal Large Language Models
S
Shi Feng
Northeastern University, China
Y
Yifei Zhang
Northeastern University, China
D
Daling Wang
Northeastern University, China