Score
Designs and builds evaluation protocols, datasets, metrics, and test suites for measuring and improving spoken‑language system robustness to accent and pronunciation variation, and implements accent identification and sampling procedures for those evaluations. Analyzes how accent variation affects model performance and develops or assesses adaptation techniques (including pronunciation evaluation) to mitigate accent-induced errors.
Current speech recognition models exhibit limited capability in phoneme-level modeling of pronunciation deviations—such as accents and disfluencies—thereby constraining the accuracy of automatic pronunciation assessment. To address this, we propose an end-to-end approach integrating multi-task learning with explicit phoneme similarity modeling, enabling fine-grained characterization of discrepancies between actual and canonical pronunciations. We construct and publicly release VCTK-accent, the first synthetic dataset specifically designed for pronunciation error modeling. Additionally, we introduce two novel metrics for quantifying pronunciation divergence. Experiments demonstrate that our method significantly improves phoneme-level transcription accuracy, particularly under non-native or atypical pronunciation conditions, and exhibits enhanced robustness compared to prior approaches. Our work establishes a new benchmark for pronunciation error detection and advances the state of the art in automatic pronunciation assessment.
Multilingual speech datasets—particularly for low-resource languages—suffer from pervasive macro-level (e.g., ambiguous dialect boundaries, absence of language planning) and micro-level (e.g., grapheme–phoneme inconsistency) quality deficiencies, severely impeding ASR model training and evaluation. This paper takes Taiwanese Hokkien (nan_tw) as a case study and proposes, for the first time, a dual-track framework integrating sociolinguistic awareness and prospective language planning to embed linguistic governance directly into ASR data curation. Through cross-dataset auditing (Common Voice, FLEURS, VoxPopuli), fieldwork, dialect annotation consistency assessment, and orthographic adaptability testing, we identify significant macro-level risks in 21 of 37 languages examined. The work yields an actionable, linguistically grounded guideline for multilingual speech dataset construction, formally adopted by Hugging Face as the v2.0 community standard.
To address the weak assessment of accent similarity in text-to-speech synthesis—particularly the low subjective reliability and failure of objective metrics for underrepresented accents—this paper proposes a novel human–machine collaborative evaluation framework. First, it improves the XAB listening test by integrating listener disagreement modeling, text-guided evaluation, and stringent inter-annotator reliability filtering to establish a lightweight, high-reliability subjective paradigm. Second, it introduces phonetically grounded objective metrics: interpretable measures based on vowel formant distance and phoneme posteriorgram similarity. Experiments demonstrate that the refined subjective protocol significantly enhances statistical power—reducing required participants by 40%. The proposed objective metrics achieve strong correlation with human judgments (Spearman’s ρ > 0.82) and substantially outperform mainstream ASR-based metrics (e.g., WER) on low-resource accents. Crucially, this work provides the first systematic analysis revealing inherent biases in WER and similar metrics for accent similarity assessment.
Conventional pronunciation error detection relies either on language-specific phonetic rules or large-scale annotated data, limiting cross-lingual applicability and increasing adaptation costs for low-resource languages. Method: We propose an end-to-end, knowledge-free approach that synthesizes personalized “canonical pronunciation” speech via voice cloning, then performs frame-level acoustic comparison—using Mel-spectrograms, F0 contours, and phoneme durations—to quantify fine-grained deviations between the learner’s utterance and the cloned reference, thereby localizing and classifying mispronunciations without phoneme alignment or rule-based modeling. Contribution/Results: The method eliminates dependence on predefined linguistic resources, drastically reducing language adaptation overhead. Experiments across multiple languages demonstrate high accuracy in detecting pronunciation deviations, strong generalization capability, and effective cross-lingual transferability—establishing a novel paradigm for intelligent pronunciation instruction in low-resource language settings.
This study reveals a significant fairness bias in English speaker verification (SV) systems—not in discriminative performance, but in calibration quality—particularly for low-resource accents. To address this, we construct the first multi-accent fairness benchmark dataset based on VoxCeleb and propose a conditional-aware backend (DCAB) coupled with a data-balancing calibration framework, enabling the first accent-conditioned calibration optimization. Experiments show that our method reduces expected calibration error (ECE) by up to 62% on low-resource accent groups; after balanced training, multiple state-of-the-art SV systems exhibit over 50% reduction in cross-accent ECE variance, markedly improving calibration fairness. Our core contribution is establishing calibration bias as a critical dimension of SV fairness and introducing a scalable, condition-aware calibration paradigm.
本文提出使用发音表示和最优传输方法,解决不同口音差异测量问题,提供了一种可解释且灵活的口音比较框架。
This study addresses the performance bottleneck of accented speech recognition under extremely low-resource conditions (fewer than 10 utterances). The authors propose a novel approach that leverages large language model (LLM)-guided phoneme-level editing, combined with a small number of target-accented samples, to generate structured accented synthetic speech for fine-tuning self-supervised automatic speech recognition (ASR) models. This work is the first to introduce LLM-driven phoneme editing into ultra-low-resource accent adaptation, revealing that perturbations in phoneme space alone constitute an effective form of data augmentation. Experimental results demonstrate significant reductions in word error rate (WER) on real accented speech and consistent improvements across speakers and under extreme data scarcity.
This study addresses core challenges in accent conversion—namely, data alignment difficulties, insufficient disentanglement of representations, data scarcity, and speaker identity preservation—by systematically tracing the field’s evolution from early rule-based signal processing techniques (e.g., spectral warping and formant analysis) to modern reference-free neural voice conversion architectures. It innovatively integrates sociolinguistic perspectives with technical analysis in a problem-driven framework, clarifying task-specific constraints and requirements across diverse application scenarios while highlighting the critical trade-off between controllability and perceptual consistency. The work further reviews prevailing datasets and evaluation methodologies, ultimately proposing a forward-looking direction toward high-fidelity, identity-preserving, and controllable accent conversion, thereby offering both a theoretical framework and practical guidance for future research.
This study addresses the challenge that existing speech quality assessment models struggle to sensitively detect localized pitch accent errors in pitch-accent languages such as Japanese. Focusing specifically on accent correctness—a previously underexplored aspect—the authors construct a controllable synthetic Japanese dataset with systematic accent errors. Building upon self-supervised speech representations, they propose a novel approach incorporating mora-conditional feature fusion, an auxiliary task for accent error localization, pairwise ranking loss, and speaker-invariant training. The resulting model significantly improves the accuracy of ranking accent error severity for both seen and unseen speakers, demonstrating strong alignment with human judgments of accent correctness.
This study addresses the entanglement of speaker identity and accent in reference audio for zero-shot text-to-speech (TTS) synthesis by proposing a training-free decoupled control method. Building upon the F5-TTS architecture, the approach integrates LoRA-based parameter-efficient fine-tuning with an exemplar encoder to disentangle speaker and accent representations. Furthermore, it enables continuous and dynamic modulation of accent intensity through guidance weights at inference time, generalizing effectively to unseen out-of-domain accents. Experimental results demonstrate that the proposed method significantly improves accent probe accuracy while preserving high speaker similarity, achieving performance comparable to that of cascaded models.