Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过使用听众面部反应来规划对话中的情感和语调,提出ReACT-TTS框架改善对话生成的自然度。
📝 Abstract
Conversational speech depends on dialogue context and the listener's immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next utterance's emotion and prosody before speech realization. On a strict dyadic MELD protocol, Temporal conditioning yields higher mean macro-F1 and VAD concordance than Text-only across ten seeds, while accuracy remains essentially unchanged. Ablations show that temporal modeling performs best among the visual variants and that an explicit early-to-late difference is unnecessary; correct listener reactions also outperform cyclic mismatches on average. In a contextual-appropriateness study with 20 speech researchers, 76% of judgments prefer Temporal, 9% Text-only, and 15% report no preference. We further connect the predicted response style to a Grad-TTS backbone for end-to-end speech realization. Overall, the results support pre-response listener dynamics as complementary cues for conversational response planning. The source code is available at https://github.com/CYJ1/ReACT-TTS_public.
Problem

Research questions and friction points this paper is trying to address.

listener facial reactions
conversational speech generation
emotion and prosody planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

listener facial reactions
response planning
Temporal conditioning
conversational speech generation
ReACT-TTS
Y
Yunji Chu
Department of Artificial Intelligence, Sogang University, Seoul, Republic of Korea