A Native-Reference Coordinate Geometry for L2 Pronunciation Deviation Using Self-Supervised Speech Models

๐Ÿ“… 2026-09-23
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บไบ†ไธ€็งๅŸบไบŽ่‡ช็›‘็ฃ่ฏญ้Ÿณๆจกๅž‹็š„ๅŽŸ็”Ÿๅ‚่€ƒๅๆ ‡ๅ‡ ไฝ•ๆ–นๆณ•๏ผŒ็”จไบŽ่ฏ„ไผฐ็ฌฌไบŒ่ฏญ่จ€ๅ‘้Ÿณๅๅทฎ๏ผŒๆ— ้œ€ๅนณ่กŒๅฝ•้Ÿณๆˆ–ไธ“้—จ็š„ๅ‘้Ÿณๆ ‡็ญพใ€‚
๐Ÿ“ Abstract
Self-supervised speech models encode rich phonetic information, but it remains unclear how to transform this information into interpretable metrics for second-language (L2) pronunciation assessment in spontaneous speech. We propose a native-reference coordinate geometry in which phone-class averages from native speech define a low-dimensional reference subspace, and L2 speech is evaluated by its distance to matching native phone-class coordinates. Unlike prior distance-based approaches, our method does not require parallel recordings with matched linguistic content or dedicated pronunciation labels. Across different self-supervised encoders and modeling choices, the resulting native-reference distances show negative Spearman correlations up to -0.5 with speaking proficiency, indicating that higher-proficiency speakers tend to lie closer to the native-reference space.
Problem

Research questions and friction points this paper is trying to address.

self-supervised speech models
phonetic information
second-language pronunciation assessment
spontaneous speech
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-supervised speech models
native-reference coordinate geometry
L2 pronunciation assessment
low-dimensional reference subspace
๐Ÿ”Ž Similar Papers
No similar papers found.