Nucleus Speculative Decoding: Plausibility-Aware Verification Beyond Exact Distribution

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low acceptance rate and limited efficiency in speculative decoding caused by conservative verification. To overcome these challenges, we propose Nucleus Speculative Decoding, which introduces a relaxed verification mechanism based on the nucleus distribution of the target model, thereby breaking through the constraints of traditional exact distribution correction. This approach is combined with a lightweight draft model parallel verification strategy and is supported by rigorous theoretical error bound analysis. Experimental results demonstrate that the proposed method achieves up to 5.16× speedup while preserving lossless generation quality, significantly outperforming standard speculative decoding. By effectively balancing inference acceleration with output fidelity, this work establishes a new paradigm for efficient large language model inference.
📝 Abstract
Speculative decoding accelerates autoregressive generation by using a lightweight draft model to propose multiple tokens that are verified by a target model in parallel. However, the standard acceptance rule focuses on exact distribution correction and rejects tokens that remain highly plausible under the target model when the draft model assigns excess probability. This conservative verification limits the number of draft tokens retained after each verification forward pass. We introduce Nucleus Speculative Decoding (NSD), a relaxed verification method that incorporates target-model plausibility into speculative decoding. NSD accepts a draft token if it satisfies the standard acceptance rule or belongs to the target model's nucleus. We theoretically characterize the distributional deviation introduced by our method and show that the single-step error is exactly determined by the draft model's excess probability within the target nucleus. We further derive sequence-level fidelity bounds that quantify how local deviations accumulate over autoregressive decoding. Experiments across multiple target models and proposal mechanisms demonstrate that NSD consistently improves speculative decoding efficiency while maintaining competitive task performance. Our method achieves throughput speedups of up to $5.16\times$ over autoregressive decoding and up to $3.15\times$ over standard speculative decoding. These improvements coincide with longer accepted lengths, allowing more output tokens to share the cost of each target verification pass. Analysis shows that plausibility-aware verification provides an effective approach for relaxed verification and speculative decoding efficiency. Our code is available at https://github.com/EIT-NLP/Nucleus-Speculative-Decoding.
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
token verification
acceptance rule
decoding efficiency
autoregressive generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Nucleus Sampling
Plausibility-Aware Verification
Autoregressive Generation
Distributional Deviation
🔎 Similar Papers
2023-10-27IACR Cryptology ePrint ArchiveCitations: 38