Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning

📅 2024-10-21
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large Audio-Language Models (LALMs) suffer from pervasive hallucination—particularly in sound event existence judgment, temporal relation reasoning, and sound source attribute identification—undermining their real-world reliability. To address this, we introduce the first systematic, multi-task benchmark for audio hallucination evaluation, comprising three fine-grained tasks: existence verification, temporal ordering, and attribute alignment. We further propose a novel multi-round chain-of-reasoning framework that integrates audio-language joint prompting with stepwise thinking mechanisms to explicitly model foundational audio semantic logic. Experiments demonstrate that our approach significantly mitigates spurious generation, temporal misalignment, and source-attribute mismatches, achieving an average accuracy improvement of 23.6% across all three tasks. This work is the first to systematically identify and enhance LALMs’ robustness at the fundamental audio semantic level.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLPKnowledge Representation and Reasoning: Common-Sense Reasoning

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Recent advancements in large audio-language models (LALMs) have shown impressive capabilities in understanding and reasoning about audio and speech information. However, these models still face challenges, including hallucinating non-existent sound events, misidentifying the order of sound events, and incorrectly attributing sound sources, which undermine their reliability and real-world application. To systematically evaluate these issues, we propose three distinct tasks: object existence, temporal order, and object attribute within audio. These tasks assess the models' comprehension of critical audio information aspects. Our experimental results reveal limitations in these fundamental tasks, underscoring the need for better models in recognizing specific sound events, determining event sequences, and identifying sound sources. To improve performance in these areas, we introduce a multi-turn chain-of-thought approach, which demonstrates significantly improved model performance across the proposed tasks.
Problem

Research questions and friction points this paper is trying to address.

Audio Language Models
Accuracy Issues
Real-world Applications
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-round Thinking Chain
Audio Language Model
Sound Event Understanding
🔎 Similar Papers
No similar papers found.