Draft in Parallel, Condition Through Depth: Adjacent Causal Injection for Speculative Decoding

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the sequence inconsistency and restricted prefix information flow inherent in parallel draft generation caused by independent token selection. We propose DSpine, a framework that introduces a full-depth adjacent causal injection mechanism, incorporating gated adjacency conditioning across all layers of the backbone network to unfold predictive features along the depth dimension for efficient parallel decoding. Furthermore, it designs a unified transmission space alongside a layer-wise output embedding supervision strategy, and optimizes system implementation through fused kernels, the SGLang serving framework, and transition caching techniques. Experimental results demonstrate that this approach significantly extends the acceptance length, achieving an average improvement of 27.8% and increasing mean throughput by 23.3% on Qwen3-8B.
📝 Abstract
Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional decoding to a lightweight module after the backbone, which limits the flow of predecessor information to successors. Our analysis of DFlash shows that early positions already form recoverable predictions in shallow layers, and that accurate adjacent predecessors help successors more when they enter earlier. We therefore propose DSpine, a drafter with causal conditioning injection throughout the backbone: at every layer, gated adjacent injection writes each predecessor's predicted feature into its successor, so the causal conditioning chain unfolds over network depth while all positions update in parallel. A unified transfer space built from the target model's output embeddings unifies layer-wise injection with predecessor-conditioned decoding, and layer-wise output-embedding supervision promotes the formation of predicted features in shallow layers. Fused kernels and a transition cache execute both efficiently in parallel within SGLang. Across seven math, code, and chat benchmarks, DSpine achieves the longest acceptance length at both temperatures on Qwen3-4B and Qwen3-8B. At temperature zero on Qwen3-8B, it raises the seven-benchmark mean from DFlash's 3.77 to 4.82 (+27.8%); in SGLang serving tests, it delivers 23.3% higher throughput than DFlash on average.
Problem

Research questions and friction points this paper is trying to address.

Speculative Decoding
Parallel Drafting
Causal Conditioning
Token Acceptance Length
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Causal Conditioning Injection
Parallel Drafting
Gated Adjacent Injection
Output-Embedding Supervision
🔎 Similar Papers
2023-12-18Neural Information Processing SystemsCitations: 52