SharpDraft: Accelerating Long-Context Speculative Decoding with Cardinality-Aware Query Scaling

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant degradation in speculative decoding speedups during long-context inference, which is caused by attention quality dilution. To mitigate this issue, we propose SharpDraft, a training-free method that introduces a novel cardinality-aware query scaling strategy to counteract attention dilution effects. By deriving a closed-form fixed-slope approximation, SharpDraft achieves precise Top-k quality correction without requiring online adaptation. Extensive evaluations across multiple benchmarks demonstrate that our approach yields 2.59× to 3.19× end-to-end speedups, outperforming both full-parameter and LoRA-based online adaptation schemes. Furthermore, SharpDraft maintains the original GPU memory footprint while substantially improving long-text generation efficiency.
📝 Abstract
Long-form reasoning makes inference expensive, and speculative decoding mitigates this cost by verifying multiple draft tokens in parallel. Its speedup, however, can fade as context grows and draft acceptance declines. We focus on attention-mass dilution: as softmax normalizes over more visible Keys, the mass concentrated on the highest-scoring Keys can decrease. We introduce SharpDraft, a training-free method that counteracts this effect through cardinality-aware Query scaling, without the computational overhead of online adaptation. Under explicit assumptions, we derive an exact top-$k$ mass correction and deploy a closed-form fixed-slope approximation. Across AIME-26, GPQA-Diamond, and LongGenBench Diary, SharpDraft achieves $2.59$-$3.19\times$ geometric-mean end-to-end speedups over target-only autoregressive decoding when applied to DFlash, PARD, and EAGLE 3.1. With DFlash, it improves decoding speed and outperforms full-parameter and LoRA-based online adaptation in end-to-end speedup, while matching the unmodified drafter's reported peak allocated GPU memory.
Problem

Research questions and friction points this paper is trying to address.

Speculative Decoding
Long-Context Inference
Attention-Mass Dilution
Draft Acceptance Rate
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Attention-Mass Dilution
Cardinality-Aware Query Scaling
Training-Free
Long-Context