π€ AI Summary
This study addresses the limitation of objective audio quality assessment methods that yield only point estimates without uncertainty quantification. We propose a confidence interval prediction approach integrating a PEAQ frontend with bagging ensemble regression. By calibrating against subjective data to generate score distributions, the method achieves performance comparable to end-to-end models using lightweight features and public datasets. Experimental results demonstrate that the proposed approach maintains mean opinion score (MOS) prediction accuracy while significantly improving alignment with subjective score distributions and confidence interval coverage rates. Furthermore, it supports approximate pairwise comparisons. Overall, this work provides a reliable, uncertainty-aware framework for audio quality evaluation.
π Abstract
Objective audio quality metrics typically provide point estimates, whereas listening tests yield score distributions from which mean opinion scores (MOS), confidence intervals (CIs), and significance decisions are derived. We propose a lightweight intrusive metric that combines a PEAQ-style perceptual front-end (ITU-R BS.1387) with a bagging ensemble of regressors. Calibration with subjective data aligns ensemble outputs with listener scores. The resulting item-dependent score distributions enable uncertainty assessment, panel-size-matched CIs, and identification of less conclusive predictions. Calibration improves agreement with subjective distributions and CI coverage across all evaluated datasets while preserving MOS accuracy. Using only 11 fixed PEAQ features, low-capacity regressors, and public training data, the method performs comparably to more data-intensive end-to-end approaches. Its output can support uncertainty-aware assessment and target listening tests towards uncertain conditions. The distributions also enable approximate pairwise comparisons, but not yet reliable significance inference.