Score
Design, build, or analyze models and decoding algorithms that predict multiple future tokens per decoding step—covering non‑autoregressive, parallel, and chunked generation—so the system outputs groups of tokens (next‑k / multi‑step) rather than a single token. This work includes techniques for entropy‑ or value‑guided selection, adaptive per‑step lengths, and speculative or cached serving alignment to reduce online decoding steps while managing token dependency, coherence, and the quality/latency tradeoffs.
Next-token prediction (NTP)-based language models suffer from fundamental limitations: weak long-horizon planning, severe error accumulation, and low computational efficiency. To address these, this paper systematically surveys alternative paradigms to NTP and proposes, for the first time, a unified five-dimensional taxonomy: multi-token prediction, plan-then-generate, latent-space reasoning, continuous-generation methods, and non-Transformer architectures. By integrating techniques—including multi-step forecasting, hierarchical planning, continuous latent-space modeling, diffusion/flow-matching, energy-based optimization, and novel neural structures—the work characterizes performance boundaries and synergistic potential across approaches. Crucially, it establishes the first comprehensive, dimensionally explicit methodology for NTP alternatives. This framework provides both theoretical foundations and concrete technical pathways toward developing efficient, controllable, and high-fidelity text generation models.
Autoregressive (AR) text generation in large language models (LLMs) suffers from slow inference due to sequential token prediction. Method: This paper systematically surveys and restructures the parallel text generation landscape, proposing the first unified taxonomy encompassing AR parallel decoding, non-autoregressive (Non-AR) modeling, diffusion-based language models, and knowledge distillation. Through theoretical analysis and empirical benchmarking across standard datasets and real-world scenarios, it characterizes the speed–quality–efficiency trade-offs inherent in each paradigm. Contribution/Results: The study identifies key acceleration pathways and synergistic integration opportunities, establishes performance boundaries, summarizes state-of-the-art advances, and highlights persistent challenges—including scalability, output consistency, and generalization. It delivers the first structured technical roadmap for efficient LLM inference, advancing the paradigm shift from serial to parallel generation.
Traditional speculative decoding (SD) enforces strict token-level distributional alignment between draft and target models, resulting in low acceptance rates and limited speedup. This work argues that task utility—e.g., code correctness or factual accuracy—is more practically meaningful than distributional fidelity, and proposes “utility alignment” as a new paradigm: only rejecting *pivot tokens*—those whose rejection demonstrably improves downstream performance. To this end, we design a lightweight classifier that dynamically identifies pivot tokens based on task-specific metrics and selectively filters non-pivot tokens during decoding. This is the first SD framework to shift the optimization objective from distribution matching to utility alignment, incorporating a pivot-aware mechanism that preserves target model accuracy while substantially increasing acceptance rates. Experiments across diverse tasks demonstrate up to 2.5× inference speedup with no degradation in task utility.
Existing speculative decoding methods treat all tokens in the draft sequence uniformly, overlooking the critical guiding role of early tokens on subsequent generation—leading to low acceptance rates and limited speedup. This work theoretically establishes, for the first time, that early tokens in the draft sequence exhibit higher predictive importance. Building on this insight, we propose a hybrid architecture: a serial Transformer head at the front end precisely models long-range dependencies among early tokens, while a lightweight parallel MLP head at the back end efficiently generates later tokens. A hierarchical computation scheduling strategy coordinates these components. Our design preserves full model compatibility while significantly improving draft quality and acceptance rate. Experiments demonstrate that our method achieves end-to-end inference speedups over state-of-the-art speculative decoding approaches across multiple mainstream LLMs, with average acceleration ratios of 1.3–1.8×.
Autoregressive decoding in large language models incurs high latency, and existing multi-token prediction methods rely on strong independence assumptions that limit modeling fidelity. Method: We propose Parallel Token Prediction (PTP), the first framework to internalize the sampling process into the model architecture, enabling joint generation of multiple semantically coherent tokens in a single Transformer forward pass while strictly preserving expressivity over any autoregressive distribution—thereby eliminating restrictive independence assumptions. PTP integrates inverse autoregressive training with explicit sampling modeling and supports both teacher-free and distillation-based training. Results: Evaluated on Vicuna-7B, PTP achieves 4.12 average accepted tokens per speculative step on Spec-Bench, maintains full modeling capability for long-sequence generation, and attains state-of-the-art performance in speculative decoding.
To address the high inference latency induced by autoregressive decoding in large language models (LLMs), this paper proposes a novel speculative decoding (SD) paradigm. The method introduces a lightweight draft model coupled with multi-token parallel sampling, enabling rapid draft generation and concurrent verification via a dedicated validation module; it further incorporates probabilistic consistency calibration to preserve output distribution fidelity. Crucially, this work achieves the first tight integration of draft generation and parallel verification—enabling 2–4× end-to-end speedup without compromising generation quality. A systematic analysis explores the SD architectural design space, validates the efficacy of verification strategies, and characterizes scalability limits. The approach is plug-and-play, fully compatible with both open-source and industrial-grade LLM deployments. By bridging theoretical insight with practical implementation, it delivers a production-ready pathway for efficient LLM inference.
Existing speculative decoding (SD) suffers from severe inference latency due to asynchronous execution between the draft and target models and fixed draft lengths, causing mutual waiting and limiting acceleration. This paper proposes PEARL, a novel framework that eliminates temporal coupling between draft and target models. PEARL introduces a synergistic pre- and post-verification mechanism to enable fully parallel draft generation and verification, and incorporates dynamic draft-length scheduling to adaptively determine the optimal speculation length per decoding step. These innovations decouple draft and target model execution both spatially and temporally. Evaluated on multiple text generation benchmarks, PEARL achieves a 4.43× speedup over autoregressive decoding and a 1.50× improvement over standard SD. The implementation is publicly available.
Existing diffusion language models typically employ fixed-depth, single-step look-ahead decoding strategies, which struggle to balance efficiency and accuracy in long-horizon generation and fail to accommodate the heterogeneity of intermediate states. This work proposes AdaLook, a novel framework that introduces, for the first time, a dynamic multi-step look-ahead mechanism guided by the variance of candidate scores. AdaLook adaptively decides whether to further unfold or expand search branches, thereby avoiding unnecessary deep computations and enabling re-initiation of look-ahead from informative intermediate states. By integrating masked diffusion language modeling with adaptive decision-making and branch expansion strategies, AdaLook substantially outperforms existing single-step approaches across multiple benchmarks, achieving comparable generation quality with significantly fewer decoding steps.
This work addresses the misalignment between the training objective of existing draft models and the inference-stage goal of maximizing consecutive token acceptance rates, which limits the acceleration performance of speculative decoding. To resolve this, the authors propose the PARD-2 framework, which reformulates the draft model’s optimization objective to prioritize overall accepted sequence length over individual token accuracy. PARD-2 introduces a Confidence-Adaptive Token (CAT) strategy, enabling a single model to uniformly support both dependent and independent target modes. By aligning the training objective with the speculative verification process and integrating a target-aligned parallel draft model with an adaptive reweighting mechanism, the method significantly enhances consecutive acceptance length. Evaluated on Llama3.1-8B, PARD-2 achieves up to 6.94× lossless speedup, outperforming EAGLE-3 and PARD by 1.9× and 1.3×, respectively.
This study addresses the inference inefficiency of large language models caused by memory bandwidth constraints and limited parallelism in single-stream autoregressive decoding. Conducting the first device-agnostic empirical investigation of speculative decoding on consumer-grade Apple Silicon hardware, the work proposes an acceleration framework that leverages a small draft model to pre-generate multiple tokens, followed by batch verification and rejection sampling using the target model. The approach strictly preserves output distribution equivalence, as validated by chi-squared tests and sequence consistency checks. Experimental results demonstrate that the optimal configuration (K=6) achieves a 1.61× speedup with a 37.8% acceptance rate, while also revealing three configurations that incur slowdowns due to pseudo-parallelization overhead or insufficient draft model quality.
Diffusion language models suffer from low inference efficiency due to the incompatibility of their bidirectional attention mechanism with KV cache reuse, necessitating a full forward pass at each denoising step. This work proposes a lightweight, training-free KV caching strategy that dynamically decides whether to recompute the cache states of the most recent k tokens by leveraging the maximum entropy of the decoding token distribution as a proxy for cache freshness. The decision overhead is constant—accounting for only 0.5% of total inference time—and independent of context length and model size. Empirical analysis further reveals prolonged post-decoding feature fluctuations across multiple steps. Evaluated on LLaDA-8B-Instruct and Dream-7B-Instruct, the method achieves 15.2–26.4× speedup on standard tasks and 22.4–24.1× on chain-of-thought tasks while maintaining competitive accuracy.
This work addresses the inefficiency of existing parallel decoding strategies in diffusion language models, which overlook the potential of early deterministic decisions to enhance global decoding efficiency. The authors propose a training-free active parallel decoding method that, for the first time, identifies and leverages a “ripple effect” during decoding: by detecting medium-entropy “pivot” positions, prospectively evaluating their impact on downstream uncertainty, and dynamically scheduling optimal decoding paths using KV cache management. Evaluated across three diffusion language models and four benchmarks spanning reasoning and code generation, the approach achieves 4–10× end-to-end speedup (up to 18× in peak cases) while preserving generation quality and consistently outperforming prior state-of-the-art baselines by up to 5.49% in accuracy across most settings.