Score
Designs, implements, and evaluates models and production inference pipelines that generate sequences by predicting each element conditioned on prior outputs; this includes specifying autoregressive architectures, decoding algorithms and their optimizations (e.g., caching, search/sampling strategies and length control), and engineering low-latency, high-throughput serving for token-by-token generation.
Autoregressive models suffer from high real-time inference latency due to sequential dependency, and conventional compression techniques—such as pruning and quantization—often incur significant accuracy degradation. To address this, we propose a unified generation–refinement decoding framework. Methodologically, we first establish a taxonomy of generation strategies—including n-gram matching and draft-model-based approaches—and refinement mechanisms—spanning single-step verification and iterative optimization. The framework integrates speculative decoding, multi-round verification, knowledge distillation from draft models, and hardware-aware scheduling for efficient deployment across heterogeneous platforms. Evaluated on text, image, and speech generation tasks, our approach achieves an average 2.1× speedup in end-to-end latency while sustaining minimal accuracy loss (<0.5% in BLEU, CLIP, and FID metrics). This work provides both scalable theoretical foundations and system-level implementation strategies for real-time large language and multimodal model applications.
A systematic analysis of large language model (LLM) inference system architectures—and the underlying technical synergies among their components—remains lacking. Method: We propose the first unified analytical framework that uncovers three foundational principles: workload forecasting, adaptive scheduling, and cost-aware compression. We introduce a deployment-paradigm-based taxonomy—categorizing systems into single-replica, multi-replica, decoupled, and serverless configurations—and comprehensively integrate key techniques including CUDA kernel optimization, continuous batching, PagedAttention, KV cache compression and persistence, weight/activation quantization, and memory offloading. Contribution/Results: Our work establishes the first holistic architecture map of LLM inference systems, explicitly characterizing inter-technique synergies and fundamental trade-offs. The resulting framework provides a systematic design guide for deploying LLMs efficiently, elastically, and cost-effectively.
To address high decoding latency and low per-task resource utilization in multi-node pipeline-parallel LLM inference, this paper proposes a pipeline-embedded dynamic speculative decoding framework. It natively integrates a draft model (LLaMA3.2-1B) into the 14-stage pipeline of a large target model (LLaMA3.1-70B), introducing a novel dynamic prediction tree mechanism that supports real-time cross-node updates and pruning—enabling tight coordination between draft token prediction and target model computation. This design significantly improves global resource utilization per task. In end-to-end decoding, it achieves 4.46×–7.79× speedup over conventional pipeline parallelism and 2.2×–2.69× over state-of-the-art tree-based speculative decoding, substantially reducing overall latency.
Autoregressive models suffer from high generation latency due to sequential, token-by-token dependency, limiting inference efficiency. This paper proposes a context-aware pipelined parallel decoding architecture: the output sequence is partitioned into multiple subsequences, each generating one new token synchronously per step, with a lightweight context synchronization mechanism ensuring coherence. The method introduces no additional parameters or KV cache overhead, preserving autoregressive modeling capability and generation quality while enabling multi-token parallelism. Experiments on question answering, summarization, and keyword generation show 1.8–2.3× speedup in inference latency, with negligible degradation (<0.5 points) in BLEU and ROUGE scores and virtually unchanged memory footprint. The core contribution is the first integration of a structured pipelined mechanism into autoregressive decoding—achieving high-quality parallel generation without any memory overhead.
Autoregressive language models generate text token-by-token, which limits parallelization and hinders efficient output of highly predictable subsequent tokens. This work proposes MARS, a method that enables instruction-tuned models to predict multiple tokens in a single forward pass through lightweight continued training, without altering the model architecture or increasing parameter count. MARS maintains compatibility with the original inference interface and allows runtime adjustment of the trade-off between generation speed and quality. It integrates block-level KV caching, confidence-threshold-based control, and multi-token prediction training. Experiments show that MARS matches or exceeds baseline performance in single-token generation while preserving accuracy in multi-token prediction, achieving 1.5–1.7× higher throughput and up to 1.71× faster inference on Qwen2.5-7B.
Existing generative paradigms—such as autoregressive and masked diffusion models—exhibit limitations in flexibility, editability, and cross-domain adaptability. To address these, this paper proposes an architecture-agnostic, abstract modeling framework for the generation process. Its core innovation is a novel generative mechanism supporting **rewritability** and **variable-length editing**, thereby overcoming constraints inherent in unidirectional prediction or fixed-step sampling. Through formal analysis of computational complexity and learnability, we theoretically establish that this paradigm achieves superior expressive power and faster convergence. Empirical evaluation demonstrates significant improvements in accuracy, controllable editing, and domain generalization across challenging tasks—including program synthesis and scientific reasoning—while offering a unified, scalable foundation for structured generation beyond natural language. (149 words)
To address the high time complexity (O(L²)) and large cache overhead (O(L²)) of autoregressive generation in convolutional sequence models, this paper proposes FutureFill. FutureFill is the first method to achieve near-linear generation complexity (O(L log L)) for convolutional models, enabled by a causality-aware convolution kernel decomposition and an incremental state update mechanism that eliminates redundant computation. Additionally, it introduces a minimal caching strategy that compresses the prefill cache to O(L), scaling linearly with sequence length. Theoretical analysis and synthetic task evaluations demonstrate that FutureFill accelerates generation by 3.2–5.8× and reduces memory footprint by one to two orders of magnitude, significantly outperforming state-of-the-art convolutional and Transformer baselines—while preserving modeling capacity and overcoming the long-standing inference efficiency bottleneck of convolutional architectures.
This work addresses the challenge of enhancing controllability and efficiency in test-time search for autoregressive models. To this end, it proposes replacing the conventional two-dimensional grid representation with a one-dimensional coarse-to-fine ordered token structure, enabling high-quality text-to-image generation without any additional training. The approach integrates a 1D ordered tokenizer, autoregressive modeling, an image-text verifier, and multiple classical search strategies—including best-of-N sampling, beam search, and look-ahead search—to systematically investigate the interplay between token structure and search algorithms. Experimental results demonstrate that the proposed structure substantially outperforms standard grid layouts in terms of test-time scalability and guidance effectiveness, marking the first demonstration of training-free, high-fidelity text-to-image synthesis within an autoregressive framework.
This work investigates whether model ensembles within the 1–3B parameter range can enhance code generation performance through execution feedback and pipeline architectures. We construct a generate-and-refine pipeline based on small language models, incorporate an execution feedback mechanism, and employ a NEAT-inspired evolutionary algorithm to search for optimal topologies. Our experiments reveal that execution feedback is pivotal—yielding performance gains exceeding four standard deviations on HumanEval and MBPP, primarily by correcting runtime errors—whereas increased topological complexity offers no significant benefit. The refinement component’s capability outweighs the identity of the generator, and single-run evaluations tend to overestimate evolutionary improvements; early stopping proves essential to prevent performance degradation. Moreover, specialized code models consistently outperform all combinations of general-purpose models.
This work addresses the limitations of traditional autoregressive models, which rely on discrete tokenization and struggle to accurately model continuous values—often leading to functional failures in precision-sensitive tasks such as semiconductor circuit design. To overcome this, the authors propose AGDC, a unified framework that enables end-to-end autoregressive generation of hybrid discrete-continuous sequences for the first time. Built upon a Transformer architecture, AGDC integrates classification-based prediction with diffusion modeling and introduces a dynamic EOS logit adjustment mechanism alongside a length regularization loss. Evaluated on a newly curated high-precision semiconductor layout dataset, ContLayNet (334K samples), and SVG-based graphics tasks, AGDC significantly outperforms both discretization-based and fixed-structure baselines, breaking through the precision bottleneck and enabling high-fidelity, variable-length vector data generation.
This work addresses the high rollout latency of autoregressive policies caused by synchronous inference, which hinders their applicability in real-time control requiring low latency and high responsiveness. The authors propose an asynchronous inference framework based on temporal tokenization and constrained decoding, enabling strict latency control and parallel multi-trajectory decoding for the first time in autoregressive policies. While preserving action smoothness, the method significantly improves response speed without sacrificing the fast convergence and strong generalization inherent to autoregressive approaches. Evaluated in both simulation and real-world environments, the proposed framework outperforms state-of-the-art flow-matching policies of comparable scale, achieving substantially higher task completion efficiency while satisfying stringent real-time constraints.