Score
Designs, builds, or analyzes conditioning modules for film-oriented speech models that inject speaker identity information into network layers — for example by inserting speaker embeddings (x-vectors) into transformer layers or by applying feature-wise modulation to encoder activations — enabling the model to adapt representations at inference without updating core weights. Focuses on integrating speaker-conditioned pathways to make film speech-processing systems robust to speaker variability and to steer model behavior based on speaker characteristics.
This work addresses the significant performance degradation of current automatic speech recognition (ASR) systems on pathological speech caused by neurological disorders. The authors propose a parameter-efficient adaptation method that injects speaker-specific x-vector embeddings into each Transformer layer of a frozen-weight SpeechLLM encoder via feature-wise linear modulation (FiLM), enabling personalized modeling without updating the base model parameters. Evaluated on bilingual (English–Spanish) pathological speech, the approach achieves recognition performance comparable to state-of-the-art adaptation techniques while effectively preserving the model’s capability on typical speech and its generalization to spoken question-answering tasks.
This work addresses the limitation of existing speaker embedding extractors, which can only describe observed speech and cannot generate controllable speaker characteristics from natural language prompts. To bridge this gap, the authors propose the ProPS framework, which encodes textual descriptions into sentence embeddings and employs a mixture density network to model a Gaussian mixture distribution in the x-vector space, enabling semantic-driven generation of speaker embeddings. This approach is the first to produce structurally coherent and attribute-controllable speaker representations conditioned on natural language, effectively aligning linguistic semantics with speaker attribute spaces. Experimental results demonstrate that the generated x-vectors exhibit high consistency across attributes such as age, gender, accent, and prosody, with both negative log-likelihood and attribute classification accuracy confirming the validity of the learned distributions.
Understanding how self-supervised speech representation models encode speaker-specific paralinguistic attributes—such as pitch, speaking rate, and energy—remains underexplored, especially in comparison to dedicated speaker embedding models. Method: We systematically analyze intermediate-layer representations from HuBERT, WavLM, and Wav2Vec 2.0, alongside CAM++, using probe-based linear classification on quantitatively annotated acoustic features. Contribution/Results: Our study is the first to comparatively evaluate multi-dimensional interpretability across layers and models. We find that deeper SSL layers achieve superior joint discriminability across paralinguistic dimensions; CAM++ excels specifically in energy classification; and intermediate layers naturally integrate acoustic and paralinguistic information, facilitating feature disentanglement. These findings establish a hierarchical, interpretable representation foundation for speaker verification and text-to-speech synthesis.
This study investigates the implicit encoding of speaker information within the feed-forward layers of self-supervised speech Transformers. To identify speaker-sensitive neurons, we align k-means-clustered self-supervised features with i-vectors. We discover, for the first time, that feed-forward neurons implicitly encode speaker gender and broad phoneme categories. Building on this finding, we propose a speaker-relevance-aware structured pruning strategy: retaining highly speaker-correlated neurons while removing low-correlation ones. Experiments demonstrate that, under substantial parameter compression (up to 40% pruning), speaker verification and identification performance remains nearly intact—equal error rate (EER) degradation is less than 0.2%. This confirms that the identified neurons serve as critical carriers of speaker representations. Our work advances understanding of internal representational mechanisms in self-supervised speech models and provides a principled, interpretable approach to model compression.
Traditional speaker embeddings, optimized for speaker identification, excessively compress intra-speaker variability, leading to inadequate prosody and emotion modeling and reduced naturalness in speech synthesis. To address this, we propose Sub-Center Speaker Embedding (SCSE), the first approach to replace single-class centers with multiple class-specific sub-centers in embedding learning—thereby explicitly modeling speech variability while preserving identification accuracy. Our method integrates a sub-center loss function, a multi-head classification layer, and an end-to-end differentiable speech synthesis or conversion framework. Experiments on voice conversion demonstrate that SCSE improves Mean Opinion Score (MOS) by 0.4 points and increases F0 dynamic range by 23%, significantly enhancing prosodic richness and overall speech naturalness.
Real-world speech is often simultaneously degraded by multiple factors such as noise, reverberation, and nonlinear distortions, leading to significant performance degradation in existing diffusion models under such composite conditions. To address this challenge, this work proposes a layer-wise conditional injection mechanism that embeds degradation-aware features—extracted by a pretrained multitask encoder—into the diffusion model’s timestep embedding and propagates them throughout all residual blocks. This approach enables effective integration of multidimensional degradation information without altering the underlying network architecture. Experimental results demonstrate that the proposed method substantially outperforms baseline models that either inject conditions only at the input layer or operate unconditionally, achieving superior speech enhancement performance and enhanced generalization across diverse real-world composite degradation scenarios.
This study addresses the lack of a unified quantitative evaluation framework for assessing the disentanglement of speaker identity and prosody in speech content representations. To this end, we construct a generative model relying solely on a single representation to systematically compare self-supervised learning (SSL) features and supervised tokens across content, identity, and prosody dimensions. Our findings reveal that disentanglement efficacy is governed by the interplay between training objectives and information capacity, rather than being determined exclusively by supervisory signals. Furthermore, we identify two distinct representational paradigms: high-fidelity reconstruction and strong disentanglement. We demonstrate that, under constrained capacity, supervised representations can effectively isolate speaker identity. These insights provide a novel theoretical foundation for advancing speech representation learning.
This work addresses the inefficiency of traditional classifier-guided diffusion models, which require separate training of a classifier and a generative model, leading to redundancy and high computational costs. To overcome this, the authors propose an efficient approach for conditional speech generation that repurposes a pretrained speech classifier as a shared backbone network. By freezing the backbone’s parameters and training only lightweight auxiliary subnetworks, the method enables conditional generation while unifying discriminative modeling and speech synthesis within a single architecture for the first time. The framework integrates noise-conditional classification, log-Mel spectrogram space modeling, and denoising score matching, achieving high-quality speech synthesis with substantially reduced memory and computational overhead.
This study addresses the challenges of fixed modulation intensity and interaction conflicts caused by head pose variations in visual speech recognition. To this end, we propose a dynamic residual FiLM framework that incorporates a deep residual weighting mechanism to adaptively regulate the modulation strength of multi-path FiLM modules by predicting input-dependent weights, thereby effectively mitigating feature interference induced by unweighted modulation. Evaluated on the LRS2 and LRS3 benchmark datasets, the proposed framework significantly reduces phoneme error rates to 15.74% and 23.91%, respectively, outperforming existing baseline models. These results validate the effectiveness of the dynamic conditional modulation strategy in complex pose-varying scenarios.
This study investigates whether individual dimensions in the representations of self-supervised speech models (specifically WavLM) encode distinct speaker-related acoustic attributes, such as pitch, gender, intensity, noise level, and the second formant. By applying principal component analysis (PCA) to disentangle model features, the authors systematically identify independent dimensions that exhibit strong correlations with these acoustic properties, establishing for the first time a clear correspondence between specific latent dimensions and interpretable speaker characteristics. Further experiments demonstrate that manipulating these dominant dimensions enables effective control over the associated speaker attributes in speech synthesis, thereby confirming both their controllability and practical utility in downstream applications.