🤖 AI Summary
This study addresses cross-modal hallucinations in audio-visual large models caused by erroneous inter-modal interference, proposing a training-free correlated evidence decoding method. The work is the first to reveal that joint inference can undermine correct unimodal predictions, motivating a decoding strategy that selects evidence on demand. Specifically, the approach leverages pointwise mutual information to quantify the contributions of unimodal versus interactive evidence, achieving precise decoding through evidence decomposition, question-aware evidence type determination, and selective enhancement. Experimental results demonstrate that the proposed method yields accuracy improvements of up to 7% across three benchmarks, while maintaining an average time-to-first-token overhead of only 1.5× compared to standard decoding.
📝 Abstract
Audio-Visual Large Language Models (AV-LLMs) remain prone to cross-modal hallucinations, where one modality incorrectly affects predictions about another. Although contrastive decoding reduces hallucinations in vision-language models, its direct extension to AV-LLMs overlooks a key challenge: different questions require different perceptual evidence, including audio, video, or their interaction. Notably, we observe that joint audio-visual inference can weaken the prediction even when a model can recover the correct answer from a single informative modality. For example, when asked which instrument is heard, a model may correctly predict violin from the audio alone. Once a video showing a guitar is added, its confidence in violin may drop. In this paper, we introduce Relevant Evidence Decoding (RED), a training-free method that identifies question-relevant evidence and selectively strengthens its contribution. RED uses pointwise mutual information to quantify the predictive support provided by audio and video beyond the question alone. It decomposes their joint contribution into audio, video, and residual interaction components. A question-only inference pass determines the required evidence type, after which the model augments the original audio-visual prediction with the corresponding PMI contribution. Across three audio-visual hallucination benchmarks and three AV-LLMs, RED improves accuracy over standard decoding by up to 7% on CMM, 6.3% on AVHBench, and 3.8% on SVHalluc, with an average relative time to first token of 1.5x standard decoding.