To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究解决了投机解码中神经草稿与基于上下文复制之间的权衡问题,通过引入SwitchSD框架,利用模型内部信号精准识别复制意图,动态切换解码策略,提高了解码效率。
📝 Abstract
Speculative Decoding (SD) has significantly accelerated Large Language Model (LLM) inference, yet existing approaches face a fundamental tradeoff between two drafting strategies: neural drafting and context-based copying. Neural drafts (e.g., EAGLE3) provide robust performance across diverse text settings, while copy-based methods achieve higher speedups in copy-intensive regimes by generating candidates faster and exploiting long repetition spans for near-perfect speculation. We analyze existing copy-based methods and find that they are prone to accidental repetitions where surface-level n-gram overlap does not reflect a structural intent to copy, leading to false-positive triggers that ultimately degrade throughput. We introduce SwitchSD, an adaptive framework that treats copying as a latent control signal of the LLM. By training lightweight probes on the target model's internal representations, SwitchSD identifies genuine copy-intent with high precision (AUC > 0.99). This allows the system to dynamically switch between neural drafting (e.g., EAGLE) and context-based copying. Our results across Llama and Qwen families demonstrate throughput gains of up to 15% over state-of-the-art baselines like EAGLE3, effectively turning copying from a noisy heuristic into a principled, model-aware decoding regime.
Problem

Research questions and friction points this paper is trying to address.

Speculative Decoding
Large Language Model
Copying
Neural Drafting
Throughput
Innovation

Methods, ideas, or system contributions that make the work stand out.

SwitchSD
adaptive framework
internal representations
copy-intent
dynamic switching
🔎 Similar Papers
No similar papers found.