🤖 AI Summary
This study addresses the superficiality of empathetic interaction and the cascading propagation of reasoning errors in large speech models by reformulating empathetic dialogue as a structured cognitive process comprising perception, mental state reasoning, and response generation. Methodologically, this work proposes an audio-anchored attention mechanism to enhance acoustic representations, introduces step-decomposed credit assignment to precisely suppress error propagation, and employs a training paradigm combining supervised fine-tuning with reinforcement learning on a multi-stage empathy dataset. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance in perception, reasoning, and response alignment, significantly improving the interaction quality of deep emotional support.
📝 Abstract
Empathetic spoken dialogue is a sophisticated cognitive process that requires not only recognizing emotions but also inferring a user's latent mental states to provide appropriate support. However, current SpeechLLMs often treat empathy as a direct input-to-response mapping, leading to"superficially warm"but emotionally hollow interactions. In addition, since empathy relies on a multi-stage process with strong inter-step dependency, errors at any intermediate step can cascade through subsequent steps and lead to inappropriate responses, while existing training paradigms lack mechanisms to precisely localize and improve such errors. In this work, we propose EchoChat, a unified framework that reformulates empathetic spoken dialogue as a structured cognitive reasoning process integrating perception, mental-state reasoning, and response generation. To support this paradigm, we first construct EchoDialogue-400K, an acoustically rich dataset for multi-stage empathetic supervision. During the SFT stage, we strengthen acoustic grounding through proposed Acoustic-Anchored Attention (AAA). During the RL stage, we further introduce a novel stage-aware optimization objective with Step-Decomposed Credit Assignment (SDCA) to localize reasoning errors and mitigate cascaded error propagation. In addition, we introduce EchoEval, an expert-annotated benchmark for multi-dimensional empathy evaluation. Extensive experiments demonstrate that EchoChat achieves state-of-the-art performance in perception, reasoning, and response alignment. Project page: https://github.com/dingdongwang/EchoChat