Broadening Uncertainty Estimation for Audio Question Answering Across Methods, Formats, and Inputs

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the tendency of audio language models to produce high-confidence hallucinations unsupported by audio evidence during question answering. It systematically evaluates uncertainty estimation methods—including probability-based, sampling-based, self-verification, and contrastive learning approaches—combined with input ablation experiments to investigate their reliability across different question types and their correlation with audio evidence. The findings reveal that model uncertainty depends primarily on audio rather than text, establishing an effective detection baseline. Notably, the first-token probability method achieves an AUROC of 0.740 without requiring additional model calls. Furthermore, removing audio input causes a significant performance degradation of 0.101 in detection, confirming the effectiveness and predictive power of this approach for open-ended question answering.
📝 Abstract
Audio-language models can produce confident answers unsupported by the audio, motivating uncertainty estimates that identify unreliable responses. We compare probability-based, sampling-based, self-verification, evidential, and contrastive measures across four open-weight models and five audio QA benchmarks. In multiple-choice evaluation, first-token measures are strongest overall, with top-1 probability achieving a mean AUROC of .740, compared with .708 for ten-sample discrete semantic entropy, while requiring no additional model calls. Across four benchmarks, shifting from multiple-choice to open-ended evaluation lowers mean accuracy from 57.6% to 36.6%, yet uncertainty remains predictive of errors: semantic entropy, maximum token entropy, and semantic agreement achieve mean AUROCs of .697, .694, and .693, respectively. To test whether uncertainty reflects the evidence available to answer the question, we perform input ablations that remove either the audio or the question. Across top-1 confidence, entropy, and sampling-based measures, removing audio reduces error-detection AUROC by .101 on average, compared with .010 when removing the question. Together, these results establish efficient uncertainty baselines and show that uncertainty in audio-language models depends substantially more on available audio evidence than on question text.
Problem

Research questions and friction points this paper is trying to address.

Uncertainty Estimation
Audio Question Answering
Audio-Language Models
Error Detection
Input Ablation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uncertainty Estimation
Audio Question Answering
Audio-Language Models
Semantic Entropy
Input Ablation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Aaron Isidore Grace
David R. Cheriton School of Computer Science, University of Waterloo, Waterloo, ON, Canada
Weiran Wang
Weiran Wang
University of Iowa
Machine learningspeech processing