๐ค AI Summary
ๆฌๆๆๅบไบไธ็งๅบไบ่ช็็ฃ่ฏญ้ณๆจกๅ็ๅ็ๅ่ๅๆ ๅ ไฝๆนๆณ๏ผ็จไบ่ฏไผฐ็ฌฌไบ่ฏญ่จๅ้ณๅๅทฎ๏ผๆ ้ๅนณ่กๅฝ้ณๆไธ้จ็ๅ้ณๆ ็ญพใ
๐ Abstract
Self-supervised speech models encode rich phonetic information, but it remains unclear how to transform this information into interpretable metrics for second-language (L2) pronunciation assessment in spontaneous speech. We propose a native-reference coordinate geometry in which phone-class averages from native speech define a low-dimensional reference subspace, and L2 speech is evaluated by its distance to matching native phone-class coordinates. Unlike prior distance-based approaches, our method does not require parallel recordings with matched linguistic content or dedicated pronunciation labels. Across different self-supervised encoders and modeling choices, the resulting native-reference distances show negative Spearman correlations up to -0.5 with speaking proficiency, indicating that higher-proficiency speakers tend to lie closer to the native-reference space.