Where Draft Trees Lose Target Mass: Exit-Guided Speculative Decoding

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low acceptance rate in tree-based speculative decoding caused by distribution mismatches between draft trees and target models. To this end, it proposes TEV, an exact verifier, alongside a training method termed ExitTrain. The core innovation lies in revealing the exit law shared by optimal verifiers, transforming target model feedback into node-level guidance signals, and integrating parallel decision mechanisms to optimize both the construction and verification of draft trees. Experimental results demonstrate that the proposed approach increases the average output block length by 13%, reduces verification latency by 15%, and achieves a 14% improvement in end-to-end speedup compared to DDTree.
📝 Abstract
Tree-based speculative decoding verifies multiple draft continuations in one target-model pass, but finite trees built from draft scores face a fundamental draft-target mismatch. We ask whether better exact verification can increase acceptance on a fixed tree and how target feedback can improve the tree itself. Through a target-flow view, we identify a canonical exit law and prove that one plus target coverage sharply bounds the expected output-block length, including the bonus token, of any exact path verifier. All optimal verifiers share the same exit and bonus-token law, already attained by representative predraw-and-follow and sequential residual verifiers. This yields Tree Exit Verification (TEV), an exact, level-parallel procedure using one exit-node decision and one bonus-token decision. The exit law also identifies missing target probability, providing node-level feedback for Exit-Guided Draft-Tree Training (ExitTrain) on inference-time draft trees. Experiments across dialogue, code, and mathematical reasoning validate fixed-tree equivalence: ExitTrain increases average output-block length by 13%, while TEV reduces verifier-stage latency by 15%, yielding a 14% end-to-end speedup over DDTree. Our results distinguish two opportunities: better draft trees for higher acceptance and more direct verification for lower latency. Code: https://github.com/hsj576/TEV.
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
draft-target mismatch
draft tree
exact verification
inference acceleration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Tree Exit Verification
Draft-Target Mismatch
Exit-Guided Training
Inference Acceleration
🔎 Similar Papers
No similar papers found.