SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of multidimensional diagnostic information in speech quality assessment and the prohibitive cost of expert annotation by proposing a 7B-parameter diagnostic speech judge built upon audio large language models. Methodologically, leveraging limited human preference data, the model is trained via a reference-conditioned cross-lingual calibration mechanism that integrates acoustic measurement mapping, prompt engineering, supervised fine-tuning, online policy distillation, and reinforcement learning. Furthermore, this work reveals the differential impacts of distinct training signals on judging behavior and rationale generation. The resulting model demonstrates robust capabilities in dimension identification and evidence citation, significantly enhancing dimensional consistency and multilingual generalization while effectively reducing misjudgment rates and generating more precise acoustic cues.
📝 Abstract
Human speech conveys rich perceptual information, such as emotion and speaker identity, yet most automatic speech quality judges reduce it to a single naturalness score. We study diagnostic speech judges: given two candidates, a diagnostic judge decides which is better, along which perceptual dimensions (e.g., timbre, emotion, timing) they differ, and which audible cues support its decision. Learning such judges is challenging: expert annotation is costly, and simply prompting a frontier audio-language model to produce labels is unreliable: our probing reveals substantial errors and unstable instruction following. We introduce SpeechCritic, which learns a diagnostic judge in a reference-conditioned cross-lingual setting from only about 300 human-labeled comparisons. Rather than replacing the frontier model, SpeechCritic calibrates it with these labels: for each dimension, it selects the acoustic measurements that agree with human judgments, maps them to A/Tie/B probabilities, and passes these to the model as non-binding hints alongside the audio. Compared with the same model labeling without hints, this raises dimension-level agreement with humans by 6.3 points and cuts the mismatch with human Tie rates by 10.4 points. We then train a 7B judge on this supervision and find that different training signals shape different judge behaviors: SFT establishes the task, OPD transfers the teacher's dimension-level strengths and weaknesses, and RL helps most on clear-cut comparisons where human raters agree. Notably, human listeners also find that RL makes rationales cite more specific, localized acoustic cues, although it never directly rewards rationale text. Finally, we show that the pipeline is language-pair agnostic by instantiating it on both English-Japanese and English-Spanish. Together, these results demonstrate a path from limited human preferences to a diagnostic speech judge.
Problem

Research questions and friction points this paper is trying to address.

diagnostic speech judge
human preferences
speech quality assessment
audio-language model
cross-lingual
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diagnostic speech judge
Acoustic calibration
Reinforcement learning
Cross-lingual
Human preference alignment
🔎 Similar Papers
No similar papers found.