🤖 AI Summary
This study addresses the tendency of large audio-language models to hallucinate and over-rely on textual priors when processing audio-visual inputs. To mitigate this, we propose a multimodal Direct Preference Optimization (DPO) method tailored for the audio domain. Using Qwen2-Audio as the backbone, our approach systematically evaluates acoustic perturbation strategies—such as temporal reversal and frequency masking—to identify optimal configurations for constructing raw-versus-distorted audio preference pairs, thereby compelling the model to generate content grounded in authentic acoustic features. Experimental results demonstrate that the proposed method improves accuracy by 14.0% on the DCASE dataset and enhances existence verification performance on Audio Hallucination by 27.4%. These findings indicate that our approach significantly strengthens the model's temporal reasoning capabilities while effectively suppressing hallucination issues.
📝 Abstract
Large Audio Language Models (LALMs) are prone to hallucinating and over-relying on text priors when simultaneously presented with audio and text inputs. To mitigate these hallucinations, we propose utilizing the multimodal Direct Preference Optimization (mDPO) objective, which forces the model to ground its generation in the acoustic input by contrasting intact and distorted audio counterparts. We extend this preference learning framework to the audio domain by applying a variety of acoustic perturbations. Evaluating the Qwen2-Audio backbone across the DCASE 2025 Challenge and AH Existence datasets, we demonstrate that extending mDPO to LALMs significantly enhances temporal reasoning in the complex DCASE dataset, and improves performance on basic existence verification in the AH benchmark. We identify temporal reversal, frequency masking, and random noise as the most effective perturbations. Ultimately, our approach achieves an absolute accuracy improvement of 14.0% on the DCASE 2025 dataset and 27.4% on AH Existence.