Thinking in Depth, Speaking Directly: Recurrent Latent Reasoning for Paralinguistically Grounded Spoken Dialogue

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the disconnection between perception and reasoning in empathetic spoken dialogue, as well as the high latency induced by explicit chain-of-thought (CoT) prompting. To this end, it proposes LoopSLM, a framework that employs a recurrent Transformer architecture for latent-space reasoning. It introduces a novel recurrent decoding block reuse mechanism to enhance acoustic grounding and adopts a two-stage decoupled training strategy that separates reasoning from response learning, thereby eliminating reliance on explicit CoT entirely. Experimental results demonstrate that LoopSLM improves reasoning accuracy by over 20 percentage points, reduces generated tokens by 64.5%, and halves inference latency. Furthermore, it outperforms Qwen3-Omni-Thinking across most evaluation metrics, achieving efficient and direct empathetic speech interaction.
📝 Abstract
Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguistic perception and make acoustic cues more explicit in replies, yet does not ensure their effective use in response planning. We call this mismatch the perception-reasoning gap. In addition, CoT may not fully capture acoustic cues in words, and generating it adds inference latency. To address these limitations, we introduce LoopSLM, which builds on looped Transformers for latent reasoning, reusing a decoder block to refine hidden states with acoustic grounding at every pass. Its two-stage training further narrows the perception-reasoning gap by separating learning to reason from learning to respond, enabling direct inference without CoT. On EchoMind, LoopSLM improves paralinguistic understanding, reasoning, and reply quality over Qwen2.5-Omni-7B. Against the CoT-SFT baseline, LoopSLM gains over 20 points in reasoning accuracy while generating 64.5% fewer tokens at half the latency. It also outperforms Qwen3-Omni-Thinking on most empathetic reply metrics with 34x lower latency. Despite training only on dialogue data, LoopSLM improves accuracy on general audio benchmarks.
Problem

Research questions and friction points this paper is trying to address.

empathetic spoken dialogue
paralinguistic perception
perception-reasoning gap
Chain-of-Thought
inference latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Reasoning
Looped Transformers
Paralinguistic Grounding
Two-stage Training
Spoken Dialogue
🔎 Similar Papers
No similar papers found.
S
Shengbo Cai
Tsinghua University
Y
Yuxiang Wang
The Chinese University of Hong Kong, Shenzhen
J
Jingran Xie
Tsinghua University
Z
Zhisheng Zhang
Tsinghua University
Shun Lei
Shun Lei
PhD student, Tsinghua University
Speech synthesisMusic generationDance GenerationSinging Voice Synthesis
Di Cao
Di Cao
University of Electronic Science and Technology of China
power systemmachine learning
T
Teddy Sun
Tencent Hunyuan
Z
Zhiyong Wu
Tsinghua University