Score
Design and build representation-learning methods and models that separate latent factors for speaker identity and accent in speech/audio, producing distinct embeddings or subspaces that can be swapped, manipulated, or independently controlled. Analyze and evaluate these systems for operations such as accent conversion while preserving speaker identity and for robustness under limited-data or low-resource conditions.
This study addresses core challenges in accent conversion—namely, data alignment difficulties, insufficient disentanglement of representations, data scarcity, and speaker identity preservation—by systematically tracing the field’s evolution from early rule-based signal processing techniques (e.g., spectral warping and formant analysis) to modern reference-free neural voice conversion architectures. It innovatively integrates sociolinguistic perspectives with technical analysis in a problem-driven framework, clarifying task-specific constraints and requirements across diverse application scenarios while highlighting the critical trade-off between controllability and perceptual consistency. The work further reviews prevailing datasets and evaluation methodologies, ultimately proposing a forward-looking direction toward high-fidelity, identity-preserving, and controllable accent conversion, thereby offering both a theoretical framework and practical guidance for future research.
Current neural text-to-speech (TTS) systems rely on speaker embeddings to control accent, yet these embeddings conflate linguistic factors such as accent with non-linguistic attributes like timbre and emotion, resulting in poor interpretability and inadequate disentanglement. This work integrates linguistically motivated phonological rules—such as flapping, retroflexion, and vowel correspondences—into neural TTS models and conducts controlled experiments to analyze the interaction between speaker embeddings and rule-based transformations. We introduce the Phoneme Shift Rate (PSR), a novel metric that quantifies the extent to which speaker embeddings preserve or override phonological rules, thereby revealing representational entanglement between accent and speaker identity. Experimental results demonstrate that combining explicit phonological rules with speaker embeddings yields more authentic accents, while embeddings alone often attenuate rule effectiveness, confirming the coupling of accent and speaker characteristics and offering a new evaluation framework for interpretable accent control.
The encoding mechanism of accent information in Discrete Speech Representation Tokens (DSRTs) remains poorly understood, and systematic evaluation frameworks or effective modeling approaches are lacking. This work proposes a unified evaluation framework featuring an Accent ABX task and cross-accent speech conversion resynthesis experiments to systematically analyze accent representation characteristics in DSRTs. Building on these insights, the study introduces novel DSRT architectures—content-specific and content-accent joint models—to enable finer-grained accent control. The findings reveal that ASR fine-tuning substantially attenuates accent information, while naive codebook reduction fails to effectively disentangle content from accent. The proposed methods significantly outperform existing approaches in accent-controllable speech generation.
This study addresses the significant performance disparities of automatic speech recognition (ASR) systems across different accents, a phenomenon whose underlying mechanisms remain poorly understood. Focusing on the Wav2Vec2 model, the authors identify—for the first time—an 8-dimensional low-dimensional subspace in the third transformer layer that densely encodes accent-related information and exhibits a notable correlation with word error rate (r = 0.26). Through representational analysis, subspace projection, controlled perturbations, and linear attenuation experiments, they demonstrate that targeted perturbations within this subspace exacerbate performance degradation (r = 0.32). Surprisingly, simple attenuation of the subspace slightly worsens fairness, challenging the prevailing assumption that erasing accent features inherently improves ASR equity.
To address the strong coupling between speaker identity and accent characteristics in multi-speaker, multi-accent text-to-speech (TTS), this paper proposes an end-to-end disentanglement framework. Our method introduces a novel multi-scale accent modeling approach—combining global utterance-level and local phoneme-level representations—integrated with adversarial speaker disentanglement, phoneme-level accent prediction, and accent modulation modules. Crucially, it enables reference-free, accent-controllable synthesis without requiring phoneme-level accent annotations. The framework supports flexible accent switching for the same speaker while preserving individual acoustic characteristics. Evaluated on an English multi-accent dataset, our model achieves a 0.42 MOS improvement and an 18.7% increase in accent similarity over baselines. Ablation studies confirm the efficacy of each component. This work establishes a new paradigm for high-fidelity, editable multi-accent TTS systems.
Traditional speaker embeddings, optimized for speaker identification, excessively compress intra-speaker variability, leading to inadequate prosody and emotion modeling and reduced naturalness in speech synthesis. To address this, we propose Sub-Center Speaker Embedding (SCSE), the first approach to replace single-class centers with multiple class-specific sub-centers in embedding learning—thereby explicitly modeling speech variability while preserving identification accuracy. Our method integrates a sub-center loss function, a multi-head classification layer, and an end-to-end differentiable speech synthesis or conversion framework. Experiments on voice conversion demonstrate that SCSE improves Mean Opinion Score (MOS) by 0.4 points and increases F0 dynamic range by 23%, significantly enhancing prosodic richness and overall speech naturalness.
This study addresses the lack of a unified quantitative evaluation framework for assessing the disentanglement of speaker identity and prosody in speech content representations. To this end, we construct a generative model relying solely on a single representation to systematically compare self-supervised learning (SSL) features and supervised tokens across content, identity, and prosody dimensions. Our findings reveal that disentanglement efficacy is governed by the interplay between training objectives and information capacity, rather than being determined exclusively by supervisory signals. Furthermore, we identify two distinct representational paradigms: high-fidelity reconstruction and strong disentanglement. We demonstrate that, under constrained capacity, supervised representations can effectively isolate speaker identity. These insights provide a novel theoretical foundation for advancing speech representation learning.
This work addresses the entanglement of linguistic content and speaker information in general-purpose representations from speech foundation models, which limits performance on downstream tasks requiring only one of these factors. To resolve this without re-pretraining, the authors propose an interventional contrastive learning approach for post-training optimization. By constructing an interventional dataset and designing a multipart contrastive loss, the method explicitly disentangles the representation into independent subspaces for content and speaker identity. This study presents the first application of interventional contrastive learning to speech representation disentanglement, achieving significant performance gains on cross-domain speaker verification tasks. Experimental results demonstrate that the learned subspaces effectively separate the two semantic attributes, confirming the efficacy of the proposed disentanglement strategy.
This study addresses the entanglement of speaker identity and accent in reference audio for zero-shot text-to-speech (TTS) synthesis by proposing a training-free decoupled control method. Building upon the F5-TTS architecture, the approach integrates LoRA-based parameter-efficient fine-tuning with an exemplar encoder to disentangle speaker and accent representations. Furthermore, it enables continuous and dynamic modulation of accent intensity through guidance weights at inference time, generalizing effectively to unseen out-of-domain accents. Experimental results demonstrate that the proposed method significantly improves accent probe accuracy while preserving high speaker similarity, achieving performance comparable to that of cascaded models.
本文提出使用发音表示和最优传输方法,解决不同口音差异测量问题,提供了一种可解释且灵活的口音比较框架。
This study investigates whether individual dimensions in the representations of self-supervised speech models (specifically WavLM) encode distinct speaker-related acoustic attributes, such as pitch, gender, intensity, noise level, and the second formant. By applying principal component analysis (PCA) to disentangle model features, the authors systematically identify independent dimensions that exhibit strong correlations with these acoustic properties, establishing for the first time a clear correspondence between specific latent dimensions and interpretable speaker characteristics. Further experiments demonstrate that manipulating these dominant dimensions enables effective control over the associated speaker attributes in speech synthesis, thereby confirming both their controllability and practical utility in downstream applications.