🤖 AI Summary
This study addresses the inherent tension in large speech models between the high latency of explicit chain-of-thought (CoT) reasoning and the inability of fixed computational budgets to adapt to varying problem complexities. To this end, we propose AURAL, a framework that introduces a novel latent-space multi-path joint chunk-wise prediction mechanism for adaptive reasoning. Furthermore, we construct AuralReason-683K, a bilingual dataset, and integrate supervised fine-tuning with difficulty-adaptive reinforcement learning to overcome the limitations of conventional single-path optimization. Experimental results demonstrate that AURAL achieves performance comparable to CoT-RL while reducing the time-to-first-token latency by 11.8× to merely 0.1 seconds on Qwen2.5-Omni. Crucially, the proposed method dynamically allocates reasoning steps according to problem difficulty, enabling efficient and flexible inference for spoken language understanding.
📝 Abstract
Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic cues further lengthens CoT and increases latency. Latent reasoning can reduce this overhead, yet existing methods often trail CoT and remain limited by single-path supervision and reasoning budgets that do not adapt to problem difficulty. We introduce AURAL, which models a distribution over multiple plausible reasoning continuations in latent space and jointly predicts chunks of future states to reduce sequential forward passes and reasoning latency. To provide initial supervision for latent reasoning, we construct AuralReason-683K: 683K bilingual speech utterances (about 1,000 hours) with concise CoT for emotion recognition, empathetic dialogue, and general reasoning. AURAL-RL then explores beyond these traces, rewarding concise reasoning that yields high-quality answers and adapting reasoning effort to each problem. Across two backbones, AURAL-RL achieves performance comparable to CoT-RL, with larger gains over the respective supervised checkpoints on most metrics. Analysis further shows that harder questions elicit more latent reasoning steps. On Qwen2.5-Omni, it reduces time to the first answer token by 11.8x, from 1.22 to 0.10 s, versus 0.05 s for direct answering.