implement speculative decoding

Design and implement decoding algorithms that accelerate autoregressive sequence generation by using a fast proposal model to propose candidate tokens which a higher-quality target model verifies, accepts, or corrects. This includes engineering the proposal/target interface, rejection- or importance-correction rules, token-level caching and batching, sampling-distribution control, and evaluating the tradeoffs among correctness, output quality, latency, and throughput.

implementspeculativedecoding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.61
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$228K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

OPT-Tree: Speculative Decoding with Adaptive Draft Tree Structure

Feb 28, 2024
JW
Jikai Wang
🏛️ Soochow University | Huawei Group

To address the slow autoregressive decoding of large language models (LLMs) and the low acceptance rates and limited speedup of existing speculative decoding methods—caused by rigid, fixed-structure draft sequences—this paper proposes Dynamic Optimal Draft Tree Speculative Decoding. We formulate the single-step expected accepted token length as the optimization objective and adaptively construct an extensible draft tree structure, thereby overcoming the limitations of static tree designs. Our method integrates a probability-driven tree search algorithm, a lightweight autoregressive draft model, and an efficient verification mechanism to enable parallel multi-token generation with lossless acceleration. Experiments across diverse LLMs and tasks demonstrate up to 3.2× decoding speedup over standard autoregressive decoding, with an average of over 10 tokens accepted per step—significantly outperforming state-of-the-art draft strategies.

Adapting draft tree structures for speculative decodingImproving inference efficiency in autoregressive language modelsMaximizing token acceptance length per decoding step

This study addresses the inference inefficiency of large language models caused by memory bandwidth constraints and limited parallelism in single-stream autoregressive decoding. Conducting the first device-agnostic empirical investigation of speculative decoding on consumer-grade Apple Silicon hardware, the work proposes an acceleration framework that leverages a small draft model to pre-generate multiple tokens, followed by batch verification and rejection sampling using the target model. The approach strictly preserves output distribution equivalence, as validated by chi-squared tests and sequence consistency checks. Experimental results demonstrate that the optimal configuration (K=6) achieves a 1.61× speedup with a 37.8% acceptance rate, while also revealing three configurations that incur slowdowns due to pseudo-parallelization overhead or insufficient draft model quality.

autoregressive decodingconsumer hardwarelarge language models

Autoregressive language models generate text token-by-token, which limits parallelization and hinders efficient output of highly predictable subsequent tokens. This work proposes MARS, a method that enables instruction-tuned models to predict multiple tokens in a single forward pass through lightweight continued training, without altering the model architecture or increasing parameter count. MARS maintains compatibility with the original inference interface and allows runtime adjustment of the trade-off between generation speed and quality. It integrates block-level KV caching, confidence-threshold-based control, and multi-token prediction training. Experiments show that MARS matches or exceeds baseline performance in single-token generation while preserving accuracy in multi-token prediction, achieving 1.5–1.7× higher throughput and up to 1.71× faster inference on Qwen2.5-7B.

autoregressive modelslanguage modelingmulti-token generation

Reflection-Window Decoding: Text Generation with Selective Refinement

Feb 05, 2025
ZT
Zeyu Tang
🏛️ Carnegie Mellon University | Mohamed bin Zayed University of Artificial Intelligence

Autoregressive decoding in large language models (LLMs) lacks backtracking capability, leading to suboptimal sequences that deviate from the globally optimal joint probability distribution. To address this, we propose an uncertainty-aware selective refinement framework. Our method introduces token-level uncertainty estimation, integrated with a sliding reflective window and a dynamic pausing mechanism, enabling localized resampling and re-decoding during generation. The framework is plug-and-play—requiring no architectural modifications—and preserves high inference efficiency while approaching joint-probability optimality. Evaluated across diverse open-ended generation and reasoning benchmarks, it significantly improves factual consistency and fluency, yielding BLEU and ROUGE gains of 2.1–4.3 points, with inference latency overhead under 12%.

Addresses suboptimal autoregressive decodingImproves text generation in LLMsIntroduces selective refinement mechanism

Unifying Autoregressive and Diffusion-Based Sequence Generation

Apr 08, 2025
NF
Nima Fathi
🏛️ ServiceNow Research

This work addresses the inherent trade-offs among generation quality, diversity, and inference efficiency between autoregressive (AR) and diffusion-based sequence generation paradigms. To unify these frameworks, we propose position-specific noise hyperschedules that parameterize both AR and diffusion processes within a single formulation; design a hybrid token-level noising mechanism that dynamically balances absorbing-noise and uniform-noising strategies to enable error correction; and introduce KV-cache-adapted attention masking to accelerate parallel decoding. Experiments on standard language modeling benchmarks demonstrate state-of-the-art perplexity, along with significant improvements in generated sequence diversity, fidelity, and robustness—while simultaneously reducing inference latency.

Introduce hyperschedules for distinct token noise schedulesPropose hybrid noising processes to fix past mistakesUnify autoregressive and diffusion models for sequence generation

Latest Papers

What's happening recently
View more

This work addresses the absence of a unified theoretical framework for autoregressive decoding strategies in speech processing, which has led to ambiguous definitions, inconsistent taxonomies, and difficulties in fair comparison. The paper introduces, for the first time, a general formal framework that precisely specifies inclusion criteria for autoregressive search and systematically categorizes and describes decoding strategies employed in neural speech generation models. By clarifying conceptual boundaries, the framework enhances comparability and evaluation consistency across strategies, streamlines the design of decoding-centric benchmarking protocols, and enables ablation studies focused specifically on search mechanisms. Consequently, it facilitates standardized analysis of inference-stage behavior in speech generation models.

auto-regressive decodingdecoding frameworksearch strategies

This work addresses the efficiency bottleneck in speculative decoding caused by distributional mismatch between general-purpose draft models and downstream tasks. To overcome this limitation, the authors propose a task-adaptive, lightweight training and fusion strategy for draft models. Specifically, they fine-tune HASS and EAGLE-2 architectures on task-specific datasets such as MathInstruct and ShareGPT, and integrate a confidence-based routing mechanism with a merged-tree verification approach. This design significantly improves both the acceptance length of speculative tokens and task adaptability. Experimental results demonstrate that task-specialized draft models achieve superior performance on their respective benchmarks, while mixed-task training enhances robustness. Moreover, the proposed fusion strategy outperforms conventional weight averaging, delivering state-of-the-art overall results across multiple evaluation metrics.

draft modelinference-time combinationspeculative decoding

Current evaluation protocols for generative systems exhibit significant biases when assessing hybrid approaches that combine autoregressive and diffusion-based decoding. This work proposes Speculative Refinement (SpecRef), a training-free hybrid decoding method that employs entropy-guided selective masking to generate draft outputs with an autoregressive model, subsequently refined by a masked diffusion language model. Through multi-faceted evaluation—encompassing execution pass rate, exact match, and log-likelihood—we uncover critical blind spots in existing benchmarks across six datasets: code benchmarks conflate structural discovery with logical correctness, multi-stage refinement degrades performance, log-likelihood poorly correlates with generation quality, and Python post-processing disproportionately disrupts non-autoregressive models. Notably, incorporating syntactic scaffolding boosts code accuracy from near zero to over 20%, underscoring the profound impact of evaluation protocols on reported results.

autoregressive decodingbenchmarkingdiffusion decoding

This work proposes Speculative Speculative Decoding (SSD), a novel decoding paradigm that overcomes the sequential bottlenecks inherent in autoregressive generation and existing speculative decoding methods. SSD achieves the first parallelization of speculation and verification by leveraging a draft model to predict verification outcomes and proactively generate subsequent tokens, thereby eliminating the overhead of draft generation. To realize this approach, we introduce the Saguaro algorithm, which systematically addresses three core challenges in SSD through a synergistic framework integrating draft–target model collaboration, verification prediction, and preemptive speculation. Experimental results demonstrate that, on open-source inference engines, SSD achieves up to 2× speedup over optimized speculative decoding and up to 5× speedup compared to conventional autoregressive decoding.

autoregressive decodinginference accelerationparallelization

This work addresses the challenges in autoregressive speech generation, where high frame rates or high-dimensional representations often lead to distributional drift and error accumulation, while low-dimensional representations compromise reconstruction fidelity. To overcome these limitations, the authors propose a jointly optimized framework featuring a low frame rate (8 Hz) yet high-dimensional (768-dim) continuous speech representation and a streaming generation architecture. Central to this approach is Locodec, a locally conditioned codec that enhances representational interpolability and coordinate identifiability, coupled with MP-ELD—a single-token autoregressive flow-matching mechanism incorporating multi-path routing and residual classifier-free guidance to effectively mitigate error propagation. Notably, the method achieves competitive word error rates (WER) and long-term stable, high-fidelity synthesis without relying on external SSL/ASR models, pretrained language models, or post-training, while maintaining high reconstruction quality and strong single-token predictability.

autoregressive speech generationerror accumulationhigh-dimensional tokens

Hot Scholars

VT

Vithursan Thangarasa

Principal Research Scientist, Cerebras Systems
MLLLMComputer VisionSparsity
LS

Lidan Shou

Professor of Computer Science, Zhejiang University
DatabaseData & Knowledge ManagementML Systems
SZ

Sheng Zhong

Nanjing University
computer networkssecurity and privacytheory of computing