How Reliable Are Predicted MOS for Reproducing Human System-Level Preferences in Speech Enhancement?

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inconsistency between predicted mean opinion scores (MOS) and human system-level preferences in speech enhancement evaluation. To overcome the limitations of conventional correlation-based assessment, it proposes the System-level Preference Accuracy (SPA) metric, which directly evaluates whether predicted MOS values support valid system comparisons. Methodologically, this work integrates single-model evaluation, ensemble learning, and domain adaptation techniques, validated through subjective listening experiments. Results demonstrate that current state-of-the-art models still yield approximately 23% preference misjudgments, revealing that relying solely on predicted MOS can lead to unreliable comparative conclusions. Ultimately, this research establishes a more rigorous evaluation paradigm for speech enhancement systems.
📝 Abstract
We investigate whether predicted Mean Opinion Scores (MOS) can reliably support system-level comparisons of speech enhancement (SE) methods by introducing system-level preference accuracy (SPA). Although MOS prediction models are widely used to evaluate SE systems, their performance is typically assessed by correlation with human-rated MOS, which does not guarantee agreement on which system is better. SPA addresses this gap by directly evaluating whether predicted and human-rated MOS yield the same system preferences. Using SPA, we systematically evaluate three settings: single prediction models, ensembling, and domain adaptation. Through experiments, SPA varies substantially across single prediction models, from 9.4% to 76.8%. Even the best model disagrees with human judgments in approximately 23% of system comparisons. Ensembling yields only limited improvement, while domain adaptation tends to substantially improve SPA in the closed condition but brings only modest gains in the more practical open condition, where neither the target systems nor the speakers are known. These results suggest that SPA can reveal errors correlation-based evaluation alone does not expose, and that predicted MOS alone can lead to unreliable conclusions in practical SE system comparison.
Problem

Research questions and friction points this paper is trying to address.

Mean Opinion Score (MOS)
Speech Enhancement
System-Level Preference
Evaluation Reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mean Opinion Score Prediction
System-level Preference Accuracy
Speech Enhancement
Domain Adaptation
Model Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.