Score
Designs and implements systems that take speech audio and produce time-aligned phoneme transcriptions and labels, detecting substitutions, omissions, insertions, or distortions relative to expected pronunciations at the phoneme level. This includes building phoneme recognition models, phoneme-level error detectors, and mispronunciation screening pipelines that localize and flag incorrect phoneme segments in recorded speech.
Current speech recognition models exhibit limited capability in phoneme-level modeling of pronunciation deviations—such as accents and disfluencies—thereby constraining the accuracy of automatic pronunciation assessment. To address this, we propose an end-to-end approach integrating multi-task learning with explicit phoneme similarity modeling, enabling fine-grained characterization of discrepancies between actual and canonical pronunciations. We construct and publicly release VCTK-accent, the first synthetic dataset specifically designed for pronunciation error modeling. Additionally, we introduce two novel metrics for quantifying pronunciation divergence. Experiments demonstrate that our method significantly improves phoneme-level transcription accuracy, particularly under non-native or atypical pronunciation conditions, and exhibits enhanced robustness compared to prior approaches. Our work establishes a new benchmark for pronunciation error detection and advances the state of the art in automatic pronunciation assessment.
Conventional pronunciation error detection relies either on language-specific phonetic rules or large-scale annotated data, limiting cross-lingual applicability and increasing adaptation costs for low-resource languages. Method: We propose an end-to-end, knowledge-free approach that synthesizes personalized “canonical pronunciation” speech via voice cloning, then performs frame-level acoustic comparison—using Mel-spectrograms, F0 contours, and phoneme durations—to quantify fine-grained deviations between the learner’s utterance and the cloned reference, thereby localizing and classifying mispronunciations without phoneme alignment or rule-based modeling. Contribution/Results: The method eliminates dependence on predefined linguistic resources, drastically reducing language adaptation overhead. Experiments across multiple languages demonstrate high accuracy in detecting pronunciation deviations, strong generalization capability, and effective cross-lingual transferability—establishing a novel paradigm for intelligent pronunciation instruction in low-resource language settings.
This study addresses the challenge of early screening for sibilant substitution disorders in Polish-speaking children under conditions of limited clinical resources by proposing a lightweight, home-deployable conservative screening protocol. The approach integrates a wav2vec2 CTC-based phoneme recognizer with forced alignment to detect mispronunciations and employs template-driven natural language generation combined with rule-based tagging to deliver interpretable feedback to caregivers. Evaluated on 559 previously unseen child utterances, the system achieves 88.7% sequence-level matching accuracy; for the target substitution errors, it attains 72.9% precision and 61.4% recall (F1 = 0.67), with a low false alarm rate of 2.7%. A clinician-in-the-loop validation mechanism is incorporated to ensure safety margins in diagnostic decision-making.
This study addresses the underexplored issue of fairness in phoneme-based automatic speech recognition (ASR) systems across demographic dimensions such as race, age, gender, and accent. It presents the first systematic evaluation of group-level biases in two open-source IPA transcription models—WhisperIPA and ZIPA—using multilingual, demographically annotated corpora. To account for linguistically acceptable phonemic variation, the authors introduce a novel metric, Soft PER (Phoneme Error Rate), which relaxes strict phoneme matching. Experimental results demonstrate that even when accommodating such permissible variation, both models exhibit significant performance disparities across languages, genders, accents, ethnicities, and age groups. These findings underscore persistent fairness challenges in current IPA-based ASR systems and highlight the need for more equitable model development and evaluation practices.
While modern text-to-speech (TTS) systems produce natural-sounding speech, they often fail to faithfully preserve phonemic contrasts that are critical for lexical and grammatical distinctions—a shortcoming undetected by conventional evaluation metrics such as Mean Opinion Score (MOS). This work proposes the first TTS evaluation framework integrating phonological knowledge, employing a phoneme classifier trained on human speech alongside phonological feature annotations (e.g., [+ATR]), acoustic cue analysis, and cross-domain transfer techniques to conduct fine-grained audits of synthetic speech. Applied to Assamese ATR vowel harmony, the framework reveals that approximately one-third of intended [+ATR] mid vowels are erroneously realized as [-ATR]. Notably, classification accuracy based on predicted phonological labels exceeds that of transcription-based labels, uncovering a systematic phonological gap between intended linguistic targets and actual TTS output.
This work addresses the challenge of pronunciation quality assessment (PQA) in low-resource languages, where phoneme-level time-aligned annotations are typically unavailable, rendering conventional methods inapplicable. The authors propose a weakly supervised PQA framework that operates without phoneme-level alignments by leveraging multilingual automatic speech recognition (ASR) to generate word-level hypotheses. From these, a phoneme confusion network is constructed to derive phoneme posterior probabilities, which are then integrated with word-level speaking rate and duration features. A cross-attention mechanism jointly models frame-level and phoneme-level information within this framework. To the best of the authors’ knowledge, this is the first approach to achieve ASR-based PQA under strictly alignment-free conditions. Experiments on the English Speechocean762 benchmark and a low-resource Tamil dataset demonstrate performance comparable to state-of-the-art methods that rely on frame-synchronized phonetic annotations, significantly enhancing applicability in resource-constrained settings.
This study addresses the challenge of automatic speech recognition (ASR) for low-resource, phonologically complex endangered languages, where it is difficult to disentangle whether performance bottlenecks stem from data scarcity or linguistic complexity. Focusing on Archi and Rutul—two Northeast Caucasian languages—the authors construct standardized speech–text datasets and introduce a phoneme-level error analysis framework. Their findings reveal an S-shaped learning curve between recognition accuracy and phoneme frequency, suggesting that model generalization can partially overcome data sparsity. By designing language-specific phonemic vocabularies and heuristic output layer initialization for wav2vec2, and benchmarking against state-of-the-art models such as Whisper and Qwen2-Audio, they demonstrate that the optimized wav2vec2 matches or even surpasses Whisper under extremely low-resource conditions. The results indicate that recognition errors are primarily attributable to data scarcity rather than phonological complexity.
This work addresses the performance disparities of automatic speech recognition systems across speaker groups by investigating the fine-grained sources of unfairness in phoneme embeddings. We propose a novel framework that attributes group-level unfairness to two distinct types of errors in phoneme modeling: stochastic error (high variance) and systematic bias. These components are disentangled using group-specific phoneme classification probes. By integrating variance and bias metrics with domain augmentation and adversarial training, we analyze the embedding properties of self-supervised speech models. Our experiments reveal that stochastic error exerts a substantially greater impact on group fairness than systematic bias, and that existing fairness-aware fine-tuning strategies struggle to effectively mitigate this issue or alter the benefits derived from probe training.