Improving Predicted MOS Scores, Not Perceived Quality: Multi-Predictor Test-Time Optimization of Enhanced Speech

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the widely used yet under-examined practice of employing non-intrusive Mean Opinion Score (MOS) predictors for speech enhancement evaluation, questioning whether optimizing these scores genuinely improves perceptual quality. We present the first systematic investigation of test-time optimization for speech enhancement, maximizing the joint mean MOS across multiple predictors via direct signal modification. Leveraging reference-free modeling and subjective MUSHRA listening tests based on the URGENT 2026 dataset, we validate the actual perceptual effects. Results demonstrate that while optimized signals achieve significantly higher MOS predictions, neither objective metrics nor subjective listening quality improve. This work exposes a critical evaluation pitfall where metric improvement does not equate to quality enhancement, confirming that such optimization can distort system comparisons. Consequently, we advocate a new paradigm mandating distinct predictors for optimization and evaluation.
📝 Abstract
Non-intrusive MOS predictors are widely used instead of subjective listening tests to evaluate and rank speech enhancement (SE) systems. If they accurately reflect perceived quality, raising their scores should lead to higher-quality speech. We present the first comprehensive analysis of test-time optimization for the SE task, which directly modifies the enhanced signal to raise the average of multiple MOS predictor scores. On seven systems from the URGENT 2026 challenge, we find that 1)~all the optimized predicted scores increase while reference-based metrics remain nearly unchanged, 2)~a non-optimized predicted score does not increase, and 3)~a MUSHRA listening test shows no improvement in perceived quality. These findings reveal a risk that such optimization can distort evaluations, e.g., biasing comparisons of SE systems regardless of their perceived quality. We believe these findings can inform future evaluation practices: they suggest that predictors used for optimization should not be used for evaluation, and that challenges should keep the predictors used for ranking undisclosed.
Problem

Research questions and friction points this paper is trying to address.

speech enhancement
MOS predictors
test-time optimization
perceived quality
evaluation bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Optimization
Non-intrusive MOS Predictors
Speech Enhancement
Perceived Quality
Evaluation Bias
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.