AcoustiClaim: A Numeric Claim Benchmark with Instrument Ground Truth

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of objective verification standards for numerical claims generated by audio language models, which impedes the assessment of their genuine acoustic perception capabilities. To this end, this work constructs an evaluation benchmark grounded in instrument-derived ground truth, which automatically extracts numerical claims from model outputs and scores them against reference measurements. By integrating natural language processing, audio analysis, and calibrated thresholding techniques, it introduces an automated scoring mechanism alongside a refusal strategy to evaluate prediction accuracy across ten acoustic quantities. The results reveal that most models perform near random baselines. Furthermore, incorporating a refusal mechanism through decoder training significantly reduces error rates in mixed-speech scenarios. Ultimately, this research establishes the first systematic evaluation framework for assessing the numerical reliability of audio language models.
📝 Abstract
Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines the quantity, and classes each quantity by where its reference can be read. Four open-weight systems and one closed model, asked for ten quantities five ways on two corpora, fill 207 cells. Of these, 49 emit fewer than five distinct values, and eight of the 158 cells that can be ranked exceed a rank correlation of 0.3, the bar we set, three with an interval clear of it, five of them one closed model reading pitch. Error sits at or above a constant-predictor floor in every ranked cell but three. The reference decoder we train declines the five voice quantities in prose on 95% of mixtures, with nothing withheld, and states them on the clean twins, reproducing its targets' rule from audio alone. With a calibrated threshold, withholding lowers error on all ten quantities on the mixtures in the mean and on eight at every split, against at most 0.6% from a random selector. A linear baseline orders errors at least as well as ours. F0 s.d. and shimmer stay above the constant floor.
Problem

Research questions and friction points this paper is trying to address.

audio language models
numeric claim verification
acoustic quantities
benchmark
instrument ground truth
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio Language Models
Numeric Claim Benchmark
Instrument Ground Truth
Reference Decoder
Calibrated Threshold
🔎 Similar Papers
No similar papers found.