Shallow Queries, Mature Values: Depth-Asynchronous Self-Speculation for Looped Transformers

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high latency inherent in autoregressive decoding for recurrent Transformers, which stems from serial computation across multiple layers. To mitigate this, we propose Deep Asynchronous Speculative Decoding (DAS). By exploiting maturity discrepancies between queries and key-values, we observe that shallow-layer queries can effectively retrieve deep-layer prefix values. Leveraging this insight, we introduce a novel deep asynchronous self-speculation mechanism and the Mature-V primitive to deeply decouple draft generation from verification. Furthermore, DAS integrates techniques including deep asynchronous prefix reading, parallel refinement, progressive block growth, and independent full-depth verification. Extensive experiments on mathematical reasoning and code generation tasks demonstrate that our approach achieves an average throughput improvement of 4.00× to 6.96× over existing baselines.
📝 Abstract
Looped Transformers reuse a shared block across recurrent depths, making autoregressive decoding expensive because every generated token requires many sequential recurrent passes. Self-speculative decoders reduce this cost by drafting at an early depth and verifying at full depth, but typically bind draft computation to prefix representations from the same recurrent depth. We find that queries and keys approach their final-depth representations earlier than values, and controlled prefix-channel interventions show that mature values substantially improve shallow draft predictions. Motivated by this asymmetry, we introduce Depth-Asynchronous Self-Speculation (DAS), which decouples the depth of draft computation from the depth of verified-prefix representations it reads. Its Mature-V primitive lets shallow queries retrieve full-depth prefix values without additional recurrent computation. We further develop DAS-Wave, which combines depth-asynchronous prefix reads with carried parallel refinement, progressive block growth, and an independent full-depth verifier. Across four recurrent-model checkpoints and mathematics and code workloads, DAS-Wave achieves 4.00--6.96$\times$ mean throughput speedup over paired full-depth autoregressive decoding in the same inference stack. These results identify prefix-information depth as an effective design axis for recurrent self-speculation.
Problem

Research questions and friction points this paper is trying to address.

Looped Transformers
autoregressive decoding
self-speculative decoding
depth-asynchronous
inference efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Depth-Asynchronous Self-Speculation
Looped Transformers
Mature-V Primitive
Self-Speculative Decoding
DAS-Wave