JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents

📅 2026-08-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses cognitive degradation, lack of empathy, and coarse-grained voice control in existing full-duplex spoken dialogue systems by proposing a modular Thinker-Talker architecture. The framework integrates speech-text joint training to preserve foundational reasoning capabilities and introduces a Persona-Adaptive Empathetic Response mechanism that leverages nonverbal cues for fine-grained empathetic responses. Natural language instruction–driven speech synthesis and efficient real-time interaction are enabled through the Joy-Duplex full-duplex protocol and state-driven turn-taking management. Experimental results demonstrate strong performance on both text-to-text (T2T) and speech-to-text (S2T) benchmarks, achieving a user interruption response rate of 0.88 and an exceptionally low false-trigger rate from background speech, thereby significantly enhancing interaction naturalness and system robustness.
📝 Abstract
We present JoyAI-Talker, a full-duplex speech dialogue system that delivers robust foundation model capabilities while empowering empathetic interaction and voice agent intelligence. JoyAI-Talker adopts a modular Thinker-Talker architecture and further implements a unified speech-text joint training pipeline to mitigate the common "cognitive degradation" bottleneck, thereby largely preserving the model's core textual reasoning, STEM, and logical capabilities while extending them to speech-based interaction. For expressive speech synthesis, the Talker module employs a text-controllable generation paradigm that enables natural-language instructions to flexibly control vocal attributes and localized paralinguistic events, such as laughter and sighs, supporting more expressive and fine-grained speech responses. To enhance conversational empathy, we introduce the Persona-Adaptive Empathetic Response (PAER) framework. PAER employs a hierarchical cognitive pipeline to extract non-verbal speaker cues, such as gender, age, and emotional state, from raw input audio, incorporate them into the Thinker's CoT reasoning, and generate context-adaptive responses that align semantically appropriate text with fine-grained control over utterance-level expressiveness and localized paralinguistic events, including sighs, speaking rate, and volume. We further integrate Joy-Duplex, a state-driven, plug-and-play full-duplex framework that functions as an efficient gating engine for real-time turn control. Extensive evaluations show that JoyAI-Talker achieves highly competitive performance on foundational T2T and S2T benchmarks. In full-duplex evaluation, the system reaches a high response rate of 0.88 under user interruptions while maintaining an extremely low false-trigger rate under background speech, demonstrating its readiness for fluid and natural speech dialogue.
Problem

Research questions and friction points this paper is trying to address.

full-duplex speech interaction
empathetic voice agents
speech-text joint modeling
paralinguistic expression
adaptive conversational response
Innovation

Methods, ideas, or system contributions that make the work stand out.

full-duplex speech interaction
empathetic voice agent
speech-text joint training
text-controllable speech synthesis
persona-adaptive response
🔎 Similar Papers
No similar papers found.