🤖 AI Summary
This study addresses the efficiency bottlenecks in speculative decoding caused by rigid verification rules and fixed tree structures. To this end, it proposes AdaptiveSpec, a training-free framework for adaptive speculative decoding. The method introduces a novel relaxed verification mechanism based on probability ratios, alongside a dynamic tree construction strategy that incorporates model confidence. These two components operate orthogonally to jointly optimize per-step verification and tree topology dynamically. Experimental evaluations conducted using the SGLang engine with the EAGLE-3 architecture demonstrate that AdaptiveSpec achieves up to a 56% throughput improvement across multiple benchmarks while recovering 93% of lossless accuracy, thereby substantially enhancing inference efficiency.
📝 Abstract
Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted, yet typically hold two decisions fixed: (1) a strict token-match verification rule and (2) a static draft-tree shape. Prior work relaxes each in isolation under limiting assumptions: long draft chains for training-free lossy verification, and adaptive tree shaping under a fixed token budget. We introduce AdaptiveSpec, a training-free per-step speculative decoding method that adapts both decisions from internal signals already produced during decoding. A per-step margin rule promotes a mismatched draft-proposed token when the ratio of the target's probability on the drafted token to its top-1 probability exceeds a threshold with no dependence on draft length or underlying drafter architecture. A per-step tree policy adjusts the draft tree's depth, width, and node count directly from a fused signal of draft top-1 confidence and a rolling acceptance history capturing recent draft-target agreement, allowing the total draft count to vary rather than only be redistributed. The two adaptations operate on orthogonal axes and compound in effect. Implemented on the SGLang production-grade serving engine, AdaptiveSpec improves throughput over the state-of-the-art autoregressive speculative decoding method EAGLE-3 by up to 56%, recovering 93% to fully lossless task accuracy across GSM8K, MATH-500, and HumanEval on three target models (DeepSeek-R1-Distill-Llama-8B, Llama-3.1-8B-Instruct, Qwen3-8B).