🤖 AI Summary
This study addresses the absence of phoneme-level intelligibility evaluation criteria for speech anonymization systems and the insufficient correlation between existing automatic metrics and human perception. To this end, we propose a novel quantitative evaluation method based on multi-ASR model ensembles with a hard voting mechanism. By combining carrier sentence tests with crowdsourced listening evaluations, this work reveals the failure mechanism of posterior probability metrics caused by miscalibration and validates the effectiveness of the proposed ensemble approach. Experimental results demonstrate that the proposed metric achieves a correlation exceeding 0.9 with human ratings under aggregated conditions, significantly enhancing evaluation accuracy. All evaluation toolchains, datasets, and source code have been made publicly available to facilitate future research.
📝 Abstract
We present the first phoneme-level intelligibility evaluation of speech anonymizers, assessing the performance of ASR-ensemble-based metrics against measured intelligibility from a crowdsourced listening test. Our results show that simple hard-voting ASR metric reaches correlations above 0.9 with human ratings when aggregated by feature, test-type, or condition, provided that multiple ASR models are combined; evaluating stimuli with and without a carrier sentence further improves the correlation at the stimulus level. However, posterior-probability-based confidence metrics bring no gain, which can be traced back to the insufficient calibration of the state-of-the-art open ASR models that were utilized here. All data, code, and evaluation tools are released as open source.