CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency of candidate token discarding and verification in block draft models for speculative decoding by proposing a tree-parallel verification framework. The method organizes scored candidate tokens into a tree structure, enabling parallel verification within a single forward pass of the target model. Furthermore, it introduces a cost-aware mechanism that adaptively determines the tree width based on latency measurements to balance verification gains against computational overhead. Notably, this approach requires no modifications to model weights or decoding rules, strictly preserving the original output distribution. Experiments across multiple domains and hardware configurations demonstrate up to 43% speedup over standard chain-based methods, significantly outperforming fixed-width strategies. The source code is publicly available.
📝 Abstract
Speculative decoding accelerates large language model inference by drafting future tokens cheaply and verifying them with the target model in parallel. Block drafters score a whole block of future tokens in one forward pass, yet standard decoding verifies only the top-scoring chain and discards the other candidates. Because these candidates are already scored, verifying more of them adds target computation but no extra drafting. We introduce CAST (Cost-Aware Speculative Trees), which packs these candidates into a tree and verifies it in a single target pass, leaving the target model, drafter weights, and decoding rule untouched. To decide how wide the tree should be, CAST adds candidates while the expected gain from the next one outweighs the verification time it adds. The width therefore adapts to each deployment from a latency measurement, without sweeping over widths. We evaluate CAST across five domains on three GPU generations and two model families. At its predicted width, CAST is faster than the standard chain in all eight settings, by up to 43%. We also find that the best width depends strongly on the deployment. Where verification cost jumps at a kernel boundary, a 128-token tree is only 2% faster than the standard chain, whereas the tree at the predicted width is 20% faster. Furthermore, we prove that CAST leaves the target output distribution unchanged under both greedy and sampled decoding. Code is available at https://github.com/js-lee-AI/CAST.
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
block drafters
inference acceleration
large language models
tree width adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Cost-Aware Speculative Trees
Block Drafters
Adaptive Tree Width
Inference Acceleration
🔎 Similar Papers
2023-12-18Neural Information Processing SystemsCitations: 52