🤖 AI Summary
This study addresses the lack of interpretability and the reliance on paired recordings in automatic pronunciation quality assessment. We propose a novel label-free, unpaired native-reference geometric approach. Specifically, features are extracted via self-supervised representation learning, and a native reference geometric space is constructed by combining singular value decomposition with context-dependent phoneme-class averaging. Second-language speech is then projected onto this space to directly quantify pronunciation deviation. Experiments on two datasets demonstrate that the proposed deviation distance exhibits a significant negative correlation with overall fluency and pronunciation quality. This work thereby achieves interpretable pronunciation analysis without requiring paired recordings.
📝 Abstract
Automatic speaking assessment systems can provide holistic proficiency scores, but often lack interpretable measures that characterize pronunciation quality. We propose a native-reference phone-class geometry for measuring second language (L2) pronunciation deviation without requiring pronunciation labels, read-aloud prompts, or matched recordings of the same text from native and L2 speakers. Given a native speech corpus, we average frame-level self-supervised representations for each context-dependent phone-class and use singular value decomposition (SVD) to derive a compact native-reference coordinate system. For each L2 utterance, we compute the corresponding averages and project them into the native-reference space. We then demonstrate that the distances between L2 and native-reference coordinates for matched phone-classes show consistent negative correlations with holistic speaking proficiency on the Dev subset of the Speak and Improve Corpus 2025 (Spearman's $ρ\!=\!-0.53$) and with pronunciation quality on the learner subset of the English Read by Japanese Students dataset ($ρ\!=\!-0.34$). These findings suggest that the proposed geometry captures acoustic-phonetic information relevant for proficiency rating while remaining applicable to spontaneous L2 speech without matched native recordings.