AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA

πŸ“… 2026-04-23
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limitations of existing audio question-answering benchmarks, which often rely on superficial acoustic cues or dataset biases and thus fail to evaluate deep auditory reasoning capabilities. To this end, the authors introduce AUDITAβ€”a large-scale, real-world audio QA benchmark comprising human-authored, fact-based questions grounded in authentic audio clips. Designed to emphasize long-range temporal dependencies and robustness against spurious correlations, AUDITA requires models to perform deep contextual reasoning over full audio narratives. Carefully crafted distractors and questions that cannot be answered in isolation effectively mitigate shortcut learning. Furthermore, item response theory (IRT) is employed to quantify both question difficulty and model proficiency. Human performance averages 32.13% accuracy, while the best current model achieves only 8.86%, starkly revealing the profound gap in complex auditory understanding among existing systems.

Technology Category

Natural Language Processing: Question AnsweringKnowledge Representation and Reasoning: Computational Complexity of ReasoningMachine Learning: Large Multimodal Models (LMMs)

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsWeb Mining and Content Analysis: Community question answering
πŸ“ Abstract
Existing audio question answering benchmarks largely emphasize sound event classification or caption-grounded queries, often enabling models to succeed through shortcut strategies, short-duration cues, lexical priors, dataset-specific biases, or even bypassing audio via metadata and captions rather than genuine reasoning Thus, we present AUDITA (Audio Understanding from Diverse Internet Trivia Authors), a large-scale, real-world benchmark to rigorously evaluate audio reasoning beyond surface-level acoustic recognition. AUDITA comprises carefully curated, human-authored trivia questions grounded in real-world audio, designed to stress robust auditory reasoning through challenging distractors and long-range temporal dependencies, using probing queries that cannot be answered from isolated text or sound cues alone. Human average accuracy of 32.13% shows both the challenge of the task while demonstrating meaningful comprehension of the audio. In stark contrast, state of-the-art audio question answering models perform poorly, with average accuracy below 8.86%. Beyond raw accuracy, we apply Item Response Theory (IRT) to estimate latent proficiency, question difficulty, and expose systematic deficiencies of the models and data.
Problem

Research questions and friction points this paper is trying to address.

audio question answering
auditory reasoning
dataset bias
shortcut learning
temporal dependencies
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio question answering
auditory reasoning
dataset bias
Item Response Theory
human-AI comparison
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Tasnim Kabir
Tasnim Kabir
University of Maryland-- College Park
Natural Language ProcessingMachine Learning
D
Dmytro Kurdydyk
Davidson College
A
Aadi Palnitkar
University of Maryland
L
Liam Dorn
Columbia University
A
Ahmed Haj Ahmed
Haverford College
J
Jordan Lee Boyd-Graber
University of Maryland