🤖 AI Summary
Self-speculative decoding suffers from throughput imbalance under fluctuating concurrency, exhibiting low effective token rates at low concurrency and excessive drafting overhead at high concurrency. This work is the first to identify and exploit the reciprocal predictability between the draft and verify phases, proposing RecGuide, a runtime dynamic orchestration framework. By integrating diffusion-model-assisted self-speculative decoding, RecGuide overlaps drafting and verification to harness idle compute at low concurrency, while adaptively allocating request-level draft chunk sizes based on workload at high concurrency, thereby achieving Pareto frontier optimization. Experimental results demonstrate that RecGuide significantly improves system throughput across diverse concurrency levels, yielding up to 1.8× speedup over baseline methods and effectively overcoming the performance bottlenecks of static strategies.
📝 Abstract
Diffusion drafting with autoregressive (AR) verification has emerged as a promising paradigm for efficient speculative decoding. Recent self-speculation models, represented by Nemotron-Labs-Diffusion, further simplify the speculative pipeline by unifying drafting and verification within a shared backbone, while enabling longer acceptance lengths. However, the Pareto frontier between aggregate and per-request throughput remains underexplored. At low concurrency, sequential draft-verify execution requires two model forward passes per round, limiting the effective tokens per forward (TPF). By contrast, at high concurrency, longer drafts incur increasingly expensive computation, forcing individual requests to operate under constrained speculation budgets and preventing full exploitation of the full-backbone drafter. Our key observation indicates that drafting and verification exhibit reciprocal predictability. Draft logits can anticipate likely verification mismatches, while recent verification outcomes predict future drafting utility and suitable block sizes. Building on this observation, we introduce Reciprocal Guidance (RecGuide), a runtime draft-verify orchestration framework that adapts speculative decoding to varying serving loads. RecGuide exploits spare compute capacity through verification-overlapped drafting at low concurrency, while dynamically allocating request-specific draft block sizes as the workload becomes increasingly compute-intensive. Experiments across a wide range of concurrency levels demonstrate consistent throughput improvements over vanilla self-speculation, achieving up to $1.8\times$ speedup.