🤖 AI Summary
This study addresses the limitation of parallel draft models in speculative decoding, which cannot continuously generate multiple rounds of drafts prior to verification due to their dependence on target model hidden states. To overcome this, we propose a recurrent speculative decoding framework that employs recurrence-aware training, enabling the draft model to adapt to its own hidden states for unverified tokens and thereby eliminating the constraint on continuous generation. Furthermore, the method incorporates a confidence-adaptive mechanism to dynamically control the number of drafting rounds and adopts a multi-stage cumulative verification strategy for unified validation, significantly reducing the forward passes required by the target model. Extensive experiments demonstrate that this approach achieves wall-clock speedup improvements ranging from 5% to 41% across various mainstream baselines while strictly preserving lossless decoding guarantees.
📝 Abstract
Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length methods decide during decoding how many draft tokens precede a verification, but they raise the speedup only for autoregressive draft models. For a parallel draft model, drafting further requires target-model hidden states for draft tokens that have not been verified. We propose DLoop, a looped form of speculative decoding that adaptively performs multiple drafting stages before verification. DLoop continues drafting while the draft model remains confident and verifies all accumulated draft tokens together. Loop-aware training keeps the draft model reliable in the additional drafting stages by exposing it to its own hidden states for unverified draft tokens. By spending additional draft-model forward passes, DLoop reduces the number of target-model forward passes required for verification. Across diverse speculative decoding methods including EAGLE-3, DFlash, Domino, DSpark, and multi-token prediction modules, DLoop improves the wall-clock speedup by 5 to 41 percent while preserving lossless decoding. Code will be available at https://github.com/naver-ai/DLoop.