Score
Design and implement methods that extract fixed-dimensional speaker embeddings from speech signals and define loss functions (e.g., embedding-matching or speaker-verification losses) that quantify mismatch between predicted and target speaker representations. Use these losses to train or regularize models so their outputs preserve or enforce the intended speaker identity.
Traditional speech enhancement methods rely on handcrafted or pretrained feature-based losses, limiting their ability to model fine-grained signal characteristics. To address this, we propose a “model-as-loss” self-consistency training paradigm: the encoder of an end-to-end differentiable encoder-decoder network serves as a dynamic, task-aware loss function, constructing a self-supervised objective in a discriminative feature space that enforces intrinsic consistency between enhanced outputs and clean speech. Crucially, this loss is fully internal—requiring no external pretrained models (e.g., WavLM or wav2vec)—and emerges solely from the current network’s own representation. Experiments demonstrate that our approach surpasses state-of-the-art deep feature losses based on WavLM/wav2vec on standard benchmarks, yields significant improvements in subjective speech quality (e.g., PESQ, STOI, and MOS), and exhibits superior generalization both within-domain and cross-domain.
Text-dependent speaker verification (SV) suffers from performance degradation when training/registration and test utterances exhibit textual mismatch. To address this, we propose a text-adaptive framework comprising: (i) a speaker-text disentanglement network that decomposes speech representations into orthogonal speaker and text embeddings; (ii) unsupervised adaptation from text-independent to text-customized speaker embeddings using only a small amount of target-text speech data without speaker labels; and (iii) post-hoc calibration of speaker embeddings via fusion with text embeddings. Experiments on RSR2015 demonstrate substantial improvements in verification accuracy under text-mismatched conditions. Notably, our method achieves text-aware embedding adaptation without requiring any target-speaker utterances—a first in the literature. This establishes a novel paradigm for low-resource, highly generalizable text-dependent SV.
This study investigates the interpretability of speaker embedding spaces by examining their mapping relationships with conventional acoustic attributes—including phonetic features, gender, and age. Using a 10,000-speaker corpus, we systematically model correlations between embedding vectors from three state-of-the-art systems (x-vector, ECAPA-TDNN, and ResNet34) and nine interpretable acoustic parameters, employing principal component analysis and linear regression. Results show that these nine parameters collectively explain over 50% of the embedding variance—matching the explanatory power of the top-10 principal components. The embedding space robustly encodes gender information, exhibiting strong implicit gender discriminability; however, it demonstrates limited capacity to represent age-related variation. To our knowledge, this is the first work to quantitatively characterize the acoustic interpretability boundary of speaker embeddings, providing empirical grounding for interpretable speech representation learning and bias analysis, as well as concrete directions for architectural and training-level refinement.
Traditional speaker embeddings, optimized for speaker identification, excessively compress intra-speaker variability, leading to inadequate prosody and emotion modeling and reduced naturalness in speech synthesis. To address this, we propose Sub-Center Speaker Embedding (SCSE), the first approach to replace single-class centers with multiple class-specific sub-centers in embedding learning—thereby explicitly modeling speech variability while preserving identification accuracy. Our method integrates a sub-center loss function, a multi-head classification layer, and an end-to-end differentiable speech synthesis or conversion framework. Experiments on voice conversion demonstrate that SCSE improves Mean Opinion Score (MOS) by 0.4 points and increases F0 dynamic range by 23%, significantly enhancing prosodic richness and overall speech naturalness.
To address speaker representation degradation and verification performance decline caused by emotional variability, this paper proposes an emotion-robust speaker representation learning framework. Methodologically, it introduces three novel components: (1) Copy-Paste speech data augmentation, (2) a cosine similarity–constrained loss function, and (3) an energy-based masking (EM) mechanism for emotion-aware spectral suppression. The EM module dynamically attenuates emotion-discriminative frequency bands using frame-level speech energy, thereby explicitly disentangling speaker identity from emotion-related variations and enhancing representation invariance. Evaluated on standard benchmarks, the proposed approach achieves a 19.29% relative reduction in equal error rate (EER) over baseline systems. Ablation studies confirm the individual efficacy and synergistic benefits of all components. This work provides a principled, reproducible methodology for building highly robust speaker verification systems under emotional variability.
本文针对真实对话中目标说话人提取的挑战,提出了一种新的损失函数来减少训练时过多静默的影响,并探讨了注册语音与目标语音不匹配的问题。
This study addresses the misalignment between model embedding geometry and human auditory perception caused by relying solely on Equal Error Rate (EER) for voice similarity evaluation. To overcome this limitation, we propose a novel perceptual alignment criterion grounded in embedding geometry. By comparing loss functions such as AAM-Softmax through rank correlation analysis and dimensional bottleneck compression, we reveal that the high-dimensional diffusion induced by classification losses fundamentally contradicts the low-dimensional nature of human perception. Notably, we identify a strong negative correlation of −0.95 between effective dimensionality and perceptual alignment. Applying dimensional bottleneck compression significantly improves the perceptual alignment score from 0.08 to 0.74. These findings provide both a theoretical foundation and a practical framework for voice similarity assessment that is substantially more consistent with human auditory perception.
This work addresses the challenge of voice anonymization by proposing a method that effectively removes speaker identity while preserving linguistic content fidelity, without relying on complex waveform reconstruction losses or explicit speaker embeddings. The approach leverages a frozen wav2vec 2.0 encoder to extract content embeddings, which are then vector-quantized and fed into a HiFi-GAN vocoder to synthesize high-quality speech. An adversarial speaker classification branch with a gradient reversal layer is introduced to deliberately confuse identity information. Requiring only content embedding alignment and adversarial training—without waveform-level losses or speaker embedding mappings—the system achieves strong performance: a word error rate of 2.53% and an anonymization equal error rate (EER) of 13.39% on the Voice Privacy Challenge (VPC) benchmark, placing it among the top-tier systems, while also unexpectedly retaining emotional characteristics with a UAR of 43.91%, yielding clear and natural-sounding output.
This work addresses the challenge that conventional speech models struggle to jointly optimize performance and computational complexity during training due to their non-differentiable architectural parameters. To overcome this limitation, the authors propose a reparameterization method based on feature noise injection, which for the first time enables end-to-end differentiable, dynamic adjustment of model architecture during training. This approach facilitates simultaneous optimization of accuracy and FLOP/s without relying on post-hoc pruning or quantization. By integrating differentiable architecture search with standard SGD optimization, the method significantly reduces computational overhead while maintaining strong performance on both voice activity detection and audio anti-spoofing tasks. The implementation has been made publicly available.
To address the challenge of simultaneously achieving discriminability and noise robustness in speaker representation learning under noisy conditions, this paper proposes a fixed-anchor-based two-stage learning framework. In the first stage, a base model is trained on clean speech data to construct highly discriminative speaker anchor representations. In the second stage, the model is fine-tuned on noisy data, where anchor-distance regularization constrains the feature space to decouple discriminative boundary stabilization from noise-induced variation suppression. The fixed anchors serve as invariant references, effectively preventing feature drift caused by joint optimization. Extensive experiments across diverse noise conditions demonstrate that the proposed method significantly outperforms end-to-end joint-optimization baselines: it preserves strong speaker discriminability while substantially improving robustness. This work establishes a novel paradigm for noise-robust speaker verification.