SEAR: Spoofing Evidence-Grounded Audio Reasoning Benchmark for Audio Language Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing audio deepfake detection benchmarks, which evaluate only the plausibility of conclusions or rationales without verifying whether models genuinely reason from acoustic evidence. To bridge this gap, this work introduces the paradigm of evidence-grounded audio reasoning, proposing the SEAR benchmark and the BAEA agent. Through a four-task Audio Question Answering (AQA) design, a frozen-model architecture augmented with external acoustic tools, and fixed or adaptive evidence acquisition strategies, the proposed framework compels models to base their detection on verifiable acoustic evidence. Experimental results demonstrate that BAEA-Fixed substantially improves both verdict accuracy and forensic quality. Furthermore, it reveals a critical discrepancy between plausible rationales and authentic evidence, confirming that misleading evidence can severely degrade detection performance.
📝 Abstract
Audio language models (ALMs) are increasingly used for audio deepfake detection (ADD), yet existing benchmarks assess their verdicts or rationale plausibility without verifying the underlying acoustic evidence. To address this issue, we first introduce spoofing evidence-grounded audio reasoning (SEAR), a four-task AQA benchmark to evaluate ALM-based ADD through acoustic evidence identification and quantification, deepfake detection, and forensic rationale generation. We further propose a bona-fide-based acoustic evidence agent (BAEA), which equips a frozen ALM with controlled acoustic tools under \textsc{fixed} or \textsc{adaptive} evidence-acquisition policies. Experiments with six ALMs reveal a clear gap between plausible rationales and verifiable acoustic evidence reasoning, while BAEA-\textsc{Fixed} improves final verdicts and forensic rationales on both evaluation partitions. Controlled interventions further show that misleading evidence degrades both detection and grounding performance.
Problem

Research questions and friction points this paper is trying to address.

Audio Deepfake Detection
Audio Language Models
Acoustic Evidence Grounding
Benchmark Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio Language Models
Audio Deepfake Detection
Acoustic Evidence Reasoning
Evidence Agent
Benchmark
🔎 Similar Papers
2024-04-22arXiv.orgCitations: 25
💼 Related Jobs
No related jobs found.
R
Rong Wan
University of Surrey
S
Suliu Qin
Singapore University of Technology and Design
J
Jiaxi Li
University of Surrey
W
Wei Xie
Guangxi University
Wenwu Wang
Wenwu Wang
Professor, University of Surrey, UK
signal processingmachine learningmachine listeningaudio/speech/audio-visualmultimodal fusion
Xiaolong Han
Xiaolong Han
University of Surrey
Deep LearningEvolutionary ComputationNeural Architecture Search
L
Lu Yin
Shenzhen University of Advanced Technology
X
Xilu Wang
University of Surrey