JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis

📅 2025-01-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the low naturalness and emotional inconsistency of synthetic speech caused by scarcity of emotional conversational speech data, this paper proposes an emotion-aware conversational text-to-speech (TTS) synthesis method. Methodologically: (1) we design a novel emotion-aware Q-former encoder to achieve fine-grained cross-modal alignment between acoustic emotional cues and textual semantics; (2) we construct a multi-module LoRA-based collaborative fine-tuning framework that jointly models emotion recognition, conversational context reasoning, and speech generation; (3) we adopt a staged training paradigm to alleviate the bottleneck imposed by limited emotional speech samples. Experiments demonstrate that our approach significantly outperforms existing methods in naturalness, emotional consistency, and dialogue-scenario adaptability, substantially improving robustness and expressiveness of emotional TTS under few-shot conditions.

Technology Category

Natural Language Processing: Conversational AI/Dialog SystemsMachine Learning: Large Multimodal Models (LMMs)Humans and AI: Emotional Intelligence

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchUser Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systems
📝 Abstract
Recently, there has been a growing demand for conversational speech synthesis (CSS) that generates more natural speech by considering the conversational context. To address this, we introduce JELLY, a novel CSS framework that integrates emotion recognition and context reasoning for generating appropriate speech in conversation by fine-tuning a large language model (LLM) with multiple partial LoRA modules. We propose an Emotion-aware Q-former encoder, which enables the LLM to perceive emotions in speech. The encoder is trained to align speech emotions with text, utilizing datasets of emotional speech. The entire model is then fine-tuned with conversational speech data to infer emotional context for generating emotionally appropriate speech in conversation. Our experimental results demonstrate that JELLY excels in emotional context modeling, synthesizing speech that naturally aligns with conversation, while mitigating the scarcity of emotional conversational speech datasets.
Problem

Research questions and friction points this paper is trying to address.

Conversational Speech Synthesis
Emotional Understanding
Data Scarcity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Emotion-aware Speech Synthesis
Contextual Conversation Modeling
Emotion-rich Data Encoder