Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding

📅 2026-07-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the efficiency bottlenecks in speculative decoding caused by rigid verification rules and fixed tree structures. To this end, it proposes AdaptiveSpec, a training-free framework for adaptive speculative decoding. The method introduces a novel relaxed verification mechanism based on probability ratios, alongside a dynamic tree construction strategy that incorporates model confidence. These two components operate orthogonally to jointly optimize per-step verification and tree topology dynamically. Experimental evaluations conducted using the SGLang engine with the EAGLE-3 architecture demonstrate that AdaptiveSpec achieves up to a 56% throughput improvement across multiple benchmarks while recovering 93% of lossless accuracy, thereby substantially enhancing inference efficiency.
📝 Abstract
Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted, yet typically hold two decisions fixed: (1) a strict token-match verification rule and (2) a static draft-tree shape. Prior work relaxes each in isolation under limiting assumptions: long draft chains for training-free lossy verification, and adaptive tree shaping under a fixed token budget. We introduce AdaptiveSpec, a training-free per-step speculative decoding method that adapts both decisions from internal signals already produced during decoding. A per-step margin rule promotes a mismatched draft-proposed token when the ratio of the target's probability on the drafted token to its top-1 probability exceeds a threshold with no dependence on draft length or underlying drafter architecture. A per-step tree policy adjusts the draft tree's depth, width, and node count directly from a fused signal of draft top-1 confidence and a rolling acceptance history capturing recent draft-target agreement, allowing the total draft count to vary rather than only be redistributed. The two adaptations operate on orthogonal axes and compound in effect. Implemented on the SGLang production-grade serving engine, AdaptiveSpec improves throughput over the state-of-the-art autoregressive speculative decoding method EAGLE-3 by up to 56%, recovering 93% to fully lossless task accuracy across GSM8K, MATH-500, and HumanEval on three target models (DeepSeek-R1-Distill-Llama-8B, Llama-3.1-8B-Instruct, Qwen3-8B).
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
LLM inference
draft tree
verification rule
throughput
Innovation

Methods, ideas, or system contributions that make the work stand out.

speculative decoding
training-free
adaptive draft tree
margin-based verification
LLM inference acceleration
🔎 Similar Papers
No similar papers found.