LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the escalating overhead of draft models in long-context speculative decoding, which erodes efficiency gains as sequence length increases. To overcome this limitation, we propose a fixed-cost parallel draft generator that leverages a block diffusion mechanism to extract multi-scale feature views from the target model. We demonstrate that the draft model need not maintain persistent states expanding with the prefix, thereby achieving complete decoupling of decoding cost from context length and enabling stateless parallel decoding. Experimental results show that our method attains state-of-the-art end-to-end efficiency across models of varying scales and real-world serving scenarios, significantly reducing per-token latency for long texts while substantially decreasing memory footprint.
📝 Abstract
Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is unnecessary. A standalone language model must grow with its prefix because it is solely responsible for every token it produces. A drafter, by contrast, only proposes candidates; the target catches and corrects every error before any token is committed. The drafter's decoding cost can therefore be made entirely independent of the prefix length. We introduce LongSpark, a block-diffusion drafter that achieves this by extracting fixed-size, multiscale views from the target's verification pass, thereby eliminating the need for a growing persistent state. Extensive evaluations demonstrate that LongSpark achieves state-of-the-art end-to-end efficiency across multiple model scales and realistic serving conditions. Notably, it delivers the lowest time-per-output-token on long-context tasks while reducing the drafter's context state by several orders of magnitude.
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
draft model
long-context
inference efficiency
autoregressive inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Block Diffusion
Fixed-cost Drafter
Long-context Inference
Multiscale Views
🔎 Similar Papers
No similar papers found.