Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical gap in the evaluation of audio language models (ALMs) as speech evaluators: their purported reliance on paralinguistic cues—such as emotion and prosody—may be illusory. To investigate, the authors propose a counterfactual auditing framework that constructs contrastive audio samples preserving textual content while varying only paralinguistic features. By jointly analyzing native judgments and performance on a recoverability control task—and further disentangling perceptual encoding from response mapping—the framework localizes the sources of model failure. Systematic evaluation across Gemini, GPT, and open-source ALMs reveals that high contrastive accuracy often masks unreliable native judgments and that heterogeneous failure modes can coexist under similar aggregate performance. These findings underscore the necessity of fine-grained behavioral auditing beyond conventional accuracy metrics and establish a new paradigm for trustworthy evaluation of audio language models.
📝 Abstract
Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paralinguistic evidence. We introduce counterfactual audits for paralinguistic response evaluation. Each audit item holds the transcript fixed while varying affect, prosody, or the timing of an affective shift, forcing a valid judge to track the audio cue rather than lexical content or response style. We evaluate ALM judges using a native one-context judgment protocol and a contrastive recoverability control, then further decompose each item into its constituent perception and response-mapping skills. This yields useful diagnostic states that identify different sources of judge failures. Across Gemini, GPT, and open audio models, we find that contrastive success often overstates native judge reliability, and that similar aggregate accuracies can hide different failure modes. These results suggest that ALM judges should not be evaluated by accuracy alone, instead requiring thorough behavioral audits before deployment.
Problem

Research questions and friction points this paper is trying to address.

audio-language models
paralinguistic evidence
response evaluation
counterfactual audits
judge reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

counterfactual audits
paralinguistic evaluation
audio-language models
behavioral diagnostics
contrastive recoverability
🔎 Similar Papers