autoregressive decoding optimization

Designs, implements, and evaluates models and production inference pipelines that generate sequences by predicting each element conditioned on prior outputs; this includes specifying autoregressive architectures, decoding algorithms and their optimizations (e.g., caching, search/sampling strategies and length control), and engineering low-latency, high-throughput serving for token-by-token generation.

autoregressivedecodingoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

A Survey of LLM Inference Systems

Jun 27, 2025
JP
James Pan
🏛️ Tsinghua University

A systematic analysis of large language model (LLM) inference system architectures—and the underlying technical synergies among their components—remains lacking. Method: We propose the first unified analytical framework that uncovers three foundational principles: workload forecasting, adaptive scheduling, and cost-aware compression. We introduce a deployment-paradigm-based taxonomy—categorizing systems into single-replica, multi-replica, decoupled, and serverless configurations—and comprehensively integrate key techniques including CUDA kernel optimization, continuous batching, PagedAttention, KV cache compression and persistence, weight/activation quantization, and memory offloading. Contribution/Results: Our work establishes the first holistic architecture map of LLM inference systems, explicitly characterizing inter-technique synergies and fundamental trade-offs. The resulting framework provides a systematic design guide for deploying LLMs efficiently, elastically, and cost-effectively.

Analyze LLM inference systems for high performance and qualityCompare techniques for model optimization and executionExplore memory management in autoregressive generation systems

Must-Read Papers

Most classic and influential ideas
View more

To address high decoding latency and low per-task resource utilization in multi-node pipeline-parallel LLM inference, this paper proposes a pipeline-embedded dynamic speculative decoding framework. It natively integrates a draft model (LLaMA3.2-1B) into the 14-stage pipeline of a large target model (LLaMA3.1-70B), introducing a novel dynamic prediction tree mechanism that supports real-time cross-node updates and pruning—enabling tight coordination between draft token prediction and target model computation. This design significantly improves global resource utilization per task. In end-to-end decoding, it achieves 4.46×–7.79× speedup over conventional pipeline parallelism and 2.2×–2.69× over state-of-the-art tree-based speculative decoding, substantially reducing overall latency.

Enhances parallelism with dynamic speculative decodingImproves global resource utilization in pipeline deploymentsReduces decoding latency in large language models

Pipelined Decoder for Efficient Context-Aware Text Generation

Jun 29, 2025
ZH
Zixian Huang
🏛️ Nanjing University | The Ohio State University

Autoregressive models suffer from high generation latency due to sequential, token-by-token dependency, limiting inference efficiency. This paper proposes a context-aware pipelined parallel decoding architecture: the output sequence is partitioned into multiple subsequences, each generating one new token synchronously per step, with a lightweight context synchronization mechanism ensuring coherence. The method introduces no additional parameters or KV cache overhead, preserving autoregressive modeling capability and generation quality while enabling multi-token parallelism. Experiments on question answering, summarization, and keyword generation show 1.8–2.3× speedup in inference latency, with negligible degradation (<0.5 points) in BLEU and ROUGE scores and virtually unchanged memory footprint. The core contribution is the first integration of a structured pipelined mechanism into autoregressive decoding—achieving high-quality parallel generation without any memory overhead.

Autoregressive models limit text generation speed due to sequential token processingEnhancing speed without compromising quality or increasing memory usageProposing a pipelined decoder for parallel context-aware text generation

Autoregressive language models generate text token-by-token, which limits parallelization and hinders efficient output of highly predictable subsequent tokens. This work proposes MARS, a method that enables instruction-tuned models to predict multiple tokens in a single forward pass through lightweight continued training, without altering the model architecture or increasing parameter count. MARS maintains compatibility with the original inference interface and allows runtime adjustment of the trade-off between generation speed and quality. It integrates block-level KV caching, confidence-threshold-based control, and multi-token prediction training. Experiments show that MARS matches or exceeds baseline performance in single-token generation while preserving accuracy in multi-token prediction, achieving 1.5–1.7× higher throughput and up to 1.71× faster inference on Qwen2.5-7B.

autoregressive modelslanguage modelingmulti-token generation

On Powerful Ways to Generate: Autoregression, Diffusion, and Beyond

Oct 07, 2025
CY
Chenxiao Yang
🏛️ Toyota Technological Institute at Chicago | Massachusetts Institute of Technology

Existing generative paradigms—such as autoregressive and masked diffusion models—exhibit limitations in flexibility, editability, and cross-domain adaptability. To address these, this paper proposes an architecture-agnostic, abstract modeling framework for the generation process. Its core innovation is a novel generative mechanism supporting **rewritability** and **variable-length editing**, thereby overcoming constraints inherent in unidirectional prediction or fixed-step sampling. Through formal analysis of computational complexity and learnability, we theoretically establish that this paradigm achieves superior expressive power and faster convergence. Empirical evaluation demonstrates significant improvements in accuracy, controllable editing, and domain generalization across challenging tasks—including program synthesis and scientific reasoning—while offering a unified, scalable foundation for structured generation beyond natural language. (149 words)

Demonstrates advantages of rewrite-capable generation for LLMsFormally studies generation processes like autoregression and diffusionQuantifies computational hardness and learnability of generation methods

FutureFill: Fast Generation from Convolutional Sequence Models

Oct 02, 2024
NA
Naman Agarwal
🏛️ Google DeepMind

To address the high time complexity (O(L²)) and large cache overhead (O(L²)) of autoregressive generation in convolutional sequence models, this paper proposes FutureFill. FutureFill is the first method to achieve near-linear generation complexity (O(L log L)) for convolutional models, enabled by a causality-aware convolution kernel decomposition and an incremental state update mechanism that eliminates redundant computation. Additionally, it introduces a minimal caching strategy that compresses the prefill cache to O(L), scaling linearly with sequence length. Theoretical analysis and synthetic task evaluations demonstrate that FutureFill accelerates generation by 3.2–5.8× and reduces memory footprint by one to two orders of magnitude, significantly outperforming state-of-the-art convolutional and Transformer baselines—while preserving modeling capacity and overcoming the long-standing inference efficiency bottleneck of convolutional architectures.

Efficient auto-regressive generation in sequence modelsMinimize prefill cache size for prompt-based generationReduce generation time from quadratic to quasilinear

Latest Papers

What's happening recently
View more

This work addresses the challenge of enhancing controllability and efficiency in test-time search for autoregressive models. To this end, it proposes replacing the conventional two-dimensional grid representation with a one-dimensional coarse-to-fine ordered token structure, enabling high-quality text-to-image generation without any additional training. The approach integrates a 1D ordered tokenizer, autoregressive modeling, an image-text verifier, and multiple classical search strategies—including best-of-N sampling, beam search, and look-ahead search—to systematically investigate the interplay between token structure and search algorithms. Experimental results demonstrate that the proposed structure substantially outperforms standard grid layouts in terms of test-time scalability and guidance effectiveness, marking the first demonstration of training-free, high-fidelity text-to-image synthesis within an autoregressive framework.

autoregressive modelsimage generationtest-time search

This work investigates whether model ensembles within the 1–3B parameter range can enhance code generation performance through execution feedback and pipeline architectures. We construct a generate-and-refine pipeline based on small language models, incorporate an execution feedback mechanism, and employ a NEAT-inspired evolutionary algorithm to search for optimal topologies. Our experiments reveal that execution feedback is pivotal—yielding performance gains exceeding four standard deviations on HumanEval and MBPP, primarily by correcting runtime errors—whereas increased topological complexity offers no significant benefit. The refinement component’s capability outweighs the identity of the generator, and single-run evaluations tend to overestimate evolutionary improvements; early stopping proves essential to prevent performance degradation. Moreover, specialized code models consistently outperform all combinations of general-purpose models.

code generationexecution feedbackmodel composition

This work addresses the limitations of traditional autoregressive models, which rely on discrete tokenization and struggle to accurately model continuous values—often leading to functional failures in precision-sensitive tasks such as semiconductor circuit design. To overcome this, the authors propose AGDC, a unified framework that enables end-to-end autoregressive generation of hybrid discrete-continuous sequences for the first time. Built upon a Transformer architecture, AGDC integrates classification-based prediction with diffusion modeling and introduces a dynamic EOS logit adjustment mechanism alongside a length regularization loss. Evaluated on a newly curated high-precision semiconductor layout dataset, ContLayNet (334K samples), and SVG-based graphics tasks, AGDC significantly outperforms both discretization-based and fixed-structure baselines, breaking through the precision bottleneck and enabling high-fidelity, variable-length vector data generation.

autoregressive generationdiscrete-continuous sequenceshigh-precision generation

This work addresses the high rollout latency of autoregressive policies caused by synchronous inference, which hinders their applicability in real-time control requiring low latency and high responsiveness. The authors propose an asynchronous inference framework based on temporal tokenization and constrained decoding, enabling strict latency control and parallel multi-trajectory decoding for the first time in autoregressive policies. While preserving action smoothness, the method significantly improves response speed without sacrificing the fast convergence and strong generalization inherent to autoregressive approaches. Evaluated in both simulation and real-world environments, the proposed framework outperforms state-of-the-art flow-matching policies of comparable scale, achieving substantially higher task completion efficiency while satisfying stringent real-time constraints.

autoregressive policieslatency boundsreal-time execution

Hot Scholars

BS

Bruno Sudret

ETH Zurich
Uncertainty QuantificationStructural reliabilitySensitivity analysisReliability-based design
SM

Stefano Marelli

Senior Scientist, Lecturer - ETH Zurich
Uncertainty QuantificationSurrogate ModelingGlobal Sensitivity AnalysisInversion
PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
SL

Sergey Levine

UC Berkeley, Physical Intelligence
Machine LearningRoboticsReinforcement Learning
JG

Jiuxiang Gu

Adobe Research
Computer VisionNatural Language ProcessingMachine Learning