DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech

๐Ÿ“… 2025-05-26
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Cross-speaker emotional voice conversion suffers from emotion leakage and speaker identity distortion due to entanglement between emotional representations and speaker characteristics. To address this, we propose the first self-supervised distillation framework for speaker-independent emotion modeling. Our method introduces a novel cluster-driven sampling strategy coupled with information perturbation to enhance emotion-speaker disentanglement; designs an emotion clustering matching mechanism to explicitly align cross-speaker emotional semantics; and employs a dual-conditional Transformer architecture to jointly model emotion and linguistic content constraints. Crucially, the framework operates without labeled emotion annotations. Experiments demonstrate significant improvements in emotion transfer accuracy (+12.7% MOS), while effectively suppressing speaker feature leakageโ€”speaker similarity degrades by only 1.3%. The approach thus achieves concurrent gains in emotional naturalness and speaker identity fidelity.

Technology Category

Natural Language Processing: SpeechMachine Learning: Unsupervised & Self-Supervised LearningCognitive Modeling & Cognitive Systems: Affective Computing

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchUser Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systemsWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
๐Ÿ“ Abstract
Cross-speaker emotion transfer in speech synthesis relies on extracting speaker-independent emotion embeddings for accurate emotion modeling without retaining speaker traits. However, existing timbre compression methods fail to fully separate speaker and emotion characteristics, causing speaker leakage and degraded synthesis quality. To address this, we propose DiEmo-TTS, a self-supervised distillation method to minimize emotional information loss and preserve speaker identity. We introduce cluster-driven sampling and information perturbation to preserve emotion while removing irrelevant factors. To facilitate this process, we propose an emotion clustering and matching approach using emotional attribute prediction and speaker embeddings, enabling generalization to unlabeled data. Additionally, we designed a dual conditioning transformer to integrate style features better. Experimental results confirm the effectiveness of our method in learning speaker-irrelevant emotion embeddings.
Problem

Research questions and friction points this paper is trying to address.

Separate speaker and emotion traits in speech synthesis
Prevent speaker leakage in cross-speaker emotion transfer
Improve emotion embedding generalization for unlabeled data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-supervised distillation for disentangled emotion embeddings
Cluster-driven sampling to preserve emotion purity
Dual conditioning transformer for better style integration
๐Ÿ”Ž Similar Papers
No similar papers found.
D
Deok-Hyeon Cho
Department of Artificial Intelligence, Korea University, Seoul, Korea
Hyung-Seok Oh
Hyung-Seok Oh
Korea Unviersity
Speech synthesis Deep Learning
Seung-Bin Kim
Seung-Bin Kim
Department of Artificial Intelligence, Korea University, Seoul, Korea
Speech Synthesis
S
Seong-Whan Lee
Department of Artificial Intelligence, Korea University, Seoul, Korea