🤖 AI Summary
Conventional pronunciation error detection relies either on language-specific phonetic rules or large-scale annotated data, limiting cross-lingual applicability and increasing adaptation costs for low-resource languages. Method: We propose an end-to-end, knowledge-free approach that synthesizes personalized “canonical pronunciation” speech via voice cloning, then performs frame-level acoustic comparison—using Mel-spectrograms, F0 contours, and phoneme durations—to quantify fine-grained deviations between the learner’s utterance and the cloned reference, thereby localizing and classifying mispronunciations without phoneme alignment or rule-based modeling. Contribution/Results: The method eliminates dependence on predefined linguistic resources, drastically reducing language adaptation overhead. Experiments across multiple languages demonstrate high accuracy in detecting pronunciation deviations, strong generalization capability, and effective cross-lingual transferability—establishing a novel paradigm for intelligent pronunciation instruction in low-resource language settings.
📝 Abstract
This paper presents a novel approach for detecting mispronunciations by analyzing deviations between a user's original speech and their voice-cloned counterpart with corrected pronunciation. We hypothesize that regions with maximal acoustic deviation between the original and cloned utterances indicate potential mispronunciations. Our method leverages recent advances in voice cloning to generate a synthetic version of the user's voice with proper pronunciation, then performs frame-by-frame comparisons to identify problematic segments. Experimental results demonstrate the effectiveness of this approach in pinpointing specific pronunciation errors without requiring predefined phonetic rules or extensive training data for each target language.