🤖 AI Summary
This study addresses the inconsistency between predicted mean opinion scores (MOS) and human system-level preferences in speech enhancement evaluation. To overcome the limitations of conventional correlation-based assessment, it proposes the System-level Preference Accuracy (SPA) metric, which directly evaluates whether predicted MOS values support valid system comparisons. Methodologically, this work integrates single-model evaluation, ensemble learning, and domain adaptation techniques, validated through subjective listening experiments. Results demonstrate that current state-of-the-art models still yield approximately 23% preference misjudgments, revealing that relying solely on predicted MOS can lead to unreliable comparative conclusions. Ultimately, this research establishes a more rigorous evaluation paradigm for speech enhancement systems.
📝 Abstract
We investigate whether predicted Mean Opinion Scores (MOS) can reliably support system-level comparisons of speech enhancement (SE) methods by introducing system-level preference accuracy (SPA). Although MOS prediction models are widely used to evaluate SE systems, their performance is typically assessed by correlation with human-rated MOS, which does not guarantee agreement on which system is better. SPA addresses this gap by directly evaluating whether predicted and human-rated MOS yield the same system preferences. Using SPA, we systematically evaluate three settings: single prediction models, ensembling, and domain adaptation. Through experiments, SPA varies substantially across single prediction models, from 9.4% to 76.8%. Even the best model disagrees with human judgments in approximately 23% of system comparisons. Ensembling yields only limited improvement, while domain adaptation tends to substantially improve SPA in the closed condition but brings only modest gains in the more practical open condition, where neither the target systems nor the speakers are known. These results suggest that SPA can reveal errors correlation-based evaluation alone does not expose, and that predicted MOS alone can lead to unreliable conclusions in practical SE system comparison.