AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent tension in large speech models between the high latency of explicit chain-of-thought (CoT) reasoning and the inability of fixed computational budgets to adapt to varying problem complexities. To this end, we propose AURAL, a framework that introduces a novel latent-space multi-path joint chunk-wise prediction mechanism for adaptive reasoning. Furthermore, we construct AuralReason-683K, a bilingual dataset, and integrate supervised fine-tuning with difficulty-adaptive reinforcement learning to overcome the limitations of conventional single-path optimization. Experimental results demonstrate that AURAL achieves performance comparable to CoT-RL while reducing the time-to-first-token latency by 11.8× to merely 0.1 seconds on Qwen2.5-Omni. Crucially, the proposed method dynamically allocates reasoning steps according to problem difficulty, enabling efficient and flexible inference for spoken language understanding.
📝 Abstract
Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic cues further lengthens CoT and increases latency. Latent reasoning can reduce this overhead, yet existing methods often trail CoT and remain limited by single-path supervision and reasoning budgets that do not adapt to problem difficulty. We introduce AURAL, which models a distribution over multiple plausible reasoning continuations in latent space and jointly predicts chunks of future states to reduce sequential forward passes and reasoning latency. To provide initial supervision for latent reasoning, we construct AuralReason-683K: 683K bilingual speech utterances (about 1,000 hours) with concise CoT for emotion recognition, empathetic dialogue, and general reasoning. AURAL-RL then explores beyond these traces, rewarding concise reasoning that yields high-quality answers and adapting reasoning effort to each problem. Across two backbones, AURAL-RL achieves performance comparable to CoT-RL, with larger gains over the respective supervised checkpoints on most metrics. Analysis further shows that harder questions elicit more latent reasoning steps. On Qwen2.5-Omni, it reduces time to the first answer token by 11.8x, from 1.22 to 0.10 s, versus 0.05 s for direct answering.
Problem

Research questions and friction points this paper is trying to address.

Speech Language Models
Latent Reasoning
Chain-of-Thought
Response Latency
Adaptive Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Reasoning
Joint Chunk Prediction
Reinforcement Learning
Speech Language Models
Adaptive Reasoning
🔎 Similar Papers
No similar papers found.
Y
Yuxiang Wang
The Chinese University of Hong Kong, Shenzhen
K
Kunyu Feng
The Chinese University of Hong Kong, Shenzhen
Yuancheng Wang
Yuancheng Wang
The Chinese University of Hong Kong, Shenzhen
Deep LearningSpeech SynthesisMusic GenerationAudio Generation
Z
Zihang Liu
Tsinghua University
S
Shengbo Cai
The Chinese University of Hong Kong, Shenzhen
Q
Qinke Ni
The Chinese University of Hong Kong, Shenzhen
W
Wan Lin
The Chinese University of Hong Kong, Shenzhen
Tao Feng
Tao Feng
The Hong Kong University of Science and Technology
Y
Yingda shen
The Chinese University of Hong Kong, Shenzhen
M
Ming-Hao Hsu
Tencent Hunyuan
Zhixian Zhao
Zhixian Zhao
Northwestern Polytechnical University
Emotion Speech RecognitionUnderstanding and Generation
L
Liqiang Zhang
Tencent Hunyuan
T
Teddy Sun
Tencent Hunyuan
S
Steve Yves
Tencent Hunyuan
Zhizheng Wu
Zhizheng Wu
The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), Mel Lab
Spoken Language ProcessingDeepFake detectionMusic Processing