Hallucination Reduction for LLM-Based Audio Understanding via Multimodal Direct Preference Optimization
This study addresses the tendency of large audio-language models to hallucinate and over-rely on textual priors when processing audio-visual inputs. To mitigate this, we propose a multimodal Direct Preference Optimization (DPO) method tailored for the audio domain. Using Qwen2-Audio as the backbone, our approach systematically evaluates acoustic perturbation strategies—such as temporal reversal and frequency masking—to identify optimal configurations for constructing raw-versus-distorted audio preference pairs, thereby compelling the model to generate content grounded in authentic acoustic features. Experimental results demonstrate that the proposed method improves accuracy by 14.0% on the DCASE dataset and enhances existence verification performance on Audio Hallucination by 27.4%. These findings indicate that our approach significantly strengthens the model's temporal reasoning capabilities while effectively suppressing hallucination issues.