🤖 AI Summary
This work addresses the tendency of audio language models to blindly follow erroneous textual inputs even when confronted with clear and contradictory audio evidence. Through a counterfactual experimental setup that fixes the audio while removing conflicting text, the study reveals that although these models encode relevant acoustic information, it is systematically overridden by textual cues during cross-modal arbitration. To mitigate this bias, the authors propose GACL (Gated Audio Counterfactual Logic), a training-free decoding rule that leverages counterfactual reasoning to recalibrate modality trust. Evaluated under strict fidelity constraints—allowing no more than a 5-percentage-point drop in faithfulness—GACL improves normalized Area Under the Curve (nAUC) by 17.8 points over the strongest baseline. Remarkably, without any task-specific tuning, GACL generalizes to vision–language arbitration tasks, yielding gains of up to 40.5 percentage points.
📝 Abstract
Audio-language models (ALMs) often follow text that conflicts with audio, even when the audio evidence is clear. This raises a basic question: is the audio-supported answer unavailable, or is it represented but overridden by the conflicting text? We examine this question using a same-audio counterfactual that keeps the audio fixed, removes only the conflicting text, and measures the resulting shift in model preference. Across five ALMs and four conflict tasks, 64.1% of conflict samples show a sign flip: the same-audio branch prefers the audio-supported answer, whereas the joint branch prefers the text-supported answer. This pattern suggests that the relevant audio evidence is encoded but loses in arbitration. Activation patching further localizes the reversal to answer-position computation, and patching effects closely track output candidate-score differences (Spearman rho=0.93). Using this diagnostic, we propose Gated Audio Counterfactual Logit Correction (GACL), a training-free decoding rule that interpolates between joint and same-audio scores. Under a strict 5 pp faithfulness-drop budget, GACL improves nAUC by 17.8 points over the best contrastive baseline and transfers without retuning to vision-text arbitration (up to +40.5 pp).