AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing conversational speech synthesis methods, which rely on predefined emotion labels that fail to capture authentic emotional expression and suffer from redundant multimodal tokens in multi-turn dialogues that hinder contextual modeling. To overcome these challenges, the authors propose the AuEmoChat framework, which first constructs a discrete latent space of genuine emotion representations derived from large-scale emotional speech (AuEmoCodec), then introduces an emotion-guided multimodal token merging algorithm (AuEmoToMe), and further incorporates an emotion flow matching mechanism to jointly optimize emotion and acoustic modeling. Experiments on the NCSSD-EmCap dataset demonstrate that the proposed approach significantly outperforms state-of-the-art systems, achieving notable improvements in both emotional expressiveness and speech naturalness.
📝 Abstract
Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions due to limited predefined emotion label spaces (e.g., seven emotion categories), while redundant multimodal tokens in multi-turn dialogue history interfere with context understanding. To address these issues, we propose AuEmoChat, a CSS framework for authentic emotion understanding and rendering. First, we develop AuEmoCodec, which learns a discrete authentic emotion token space from large-scale emotional speech via finite scalar quantization, enabling a more authentic emotion representation than limited basic emotion categories. We further propose AuEmoToMe, an authentic-emotion-guided token merging algorithm that merges redundant tokens in multimodal dialogue history while preserving emotion-relevant context. We integrate it into an autoregressive text-speech model to predict the target authentic emotion token and speech tokens. Finally, we propose Authentic Emotion Flow Matching, which renders speech by jointly conditioning on merged dialogue context, target authentic emotion, and acoustic priors. Extensive experiments on the NCSSD-EmCap dataset demonstrate that AuEmoChat outperforms state-of-the-art CSS baselines and generates more expressive and authentic emotional speech.
Problem

Research questions and friction points this paper is trying to address.

Conversational Speech Synthesis
Authentic Emotion
Emotion Representation
Multimodal Dialogue History
Emotional Speech Synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Authentic Emotion Representation
Discrete Emotion Token Space
Emotion-Guided Token Merging
Conversational Speech Synthesis
Flow Matching
🔎 Similar Papers
No similar papers found.