Score
Designs, builds, or evaluates autoregressive generative models that reconstruct or denoise visual data by sequentially predicting pixels or discrete visual tokens; these systems produce progressive, prefix-conditioned or multiscale token maps and model global-to-local structural dependencies via next-token/next-scale probability predictions.
Visual autoregressive modeling faces challenges in scalability, long-range dependency capture, computational efficiency, and geometric/physical consistency—particularly for image, video, 3D, and multimodal generation. Method: We systematically survey ~250 works, unifying pixel-level, token-level, and scale-level representations; integrating sequence modeling, discrete representation learning (e.g., VQ-VAE, DALL·E tokenizer), causal attention, and hierarchical autoregressive decoding; and establishing connections to diffusion models and GANs. Contribution/Results: We propose the first comprehensive taxonomy spanning representation granularities and cross-cutting dimensions (hierarchical, multimodal, task-agnostic), construct an open-source knowledge base, identify core bottlenecks—including 3D structural priors and inference latency—and outline future directions in scalable architectures, efficient sampling, and physics-aware generation.
This work addresses the longstanding trade-off between generation quality and reconstruction fidelity in autoregressive image synthesis, which traditionally relies on visual tokenizers trained in a separate, multi-stage pipeline. To overcome this limitation, the authors propose the first end-to-end joint training framework that simultaneously optimizes a 1D semantic tokenizer and an autoregressive generative model. Central to their approach is the integration of a vision foundation model to enrich the semantic expressiveness of the learned tokens, along with a novel mechanism that directly supervises tokenizer learning through the generated outputs—thereby departing from conventional two-stage paradigms. Evaluated on class-unconditional ImageNet generation at 256×256 resolution, the method achieves a new state-of-the-art FID of 1.48.
To address the challenges of excessively long sequences, high training costs, and difficulty in leveraging intrinsic image hierarchies in autoregressive image generation, this paper proposes the Next Patch Prediction (NPP) paradigm. NPP aggregates low-level image tokens into high-information-density patch tokens, drastically reducing sequence length; it introduces a novel multi-scale, coarse-to-fine hierarchical patch grouping strategy that requires no model architecture modifications, additional parameters, or custom tokenizers—ensuring generality and plug-and-play compatibility. Evaluated under standard autoregressive modeling and FID assessment on ImageNet, NPP reduces training cost to approximately 0.6× that of baseline methods while improving FID by up to 1.0. The improvement is consistent across models ranging from 100M to 1.4B parameters. The core contribution is the first seamless integration of inherent image hierarchy into autoregressive generation—achieving simultaneous gains in efficiency, sample quality, and deployment practicality.
Autoregressive visual generation suffers from low inference efficiency due to sequential, token-by-token decoding. This paper proposes a dependency-aware parallelization strategy that requires no modification to the model architecture or tokenizer. It dynamically partitions the decoding process into parallel and serial regions based on the conditional dependency strength among visual tokens—marking the first integration of explicit dependency modeling into parallel decoding decisions. The method comprises dependency-aware grouping sampling and localized serial constraints, and is plug-and-play on standard Transformer decoders. Evaluated on image and video generation tasks using ImageNet and UCF-101, it achieves up to 9.5× speedup; critically, it attains 3.6× acceleration without compromising generation quality—outperforming existing parallel decoding approaches. Its core contributions are: (i) zero-modification deployment, (ii) dependency-driven parallelization, and (iii) a general, efficient, and quality-preserving acceleration paradigm for autoregressive visual generation.
Existing vision autoregressive models rely on raster-scan ordering for “next-token prediction,” neglecting the intrinsic spatial-temporal locality of visual data. This work proposes Neighborhood Autoregressive Modeling (NAR), reformulating generation as a progressive outpainting process ordered by increasing Manhattan distance from a seed region. It introduces a novel “next-neighborhood prediction” mechanism and a multi-dimensional orthogonal decoding head, enabling parallel prediction of spatial-temporal neighborhood tokens and drastically reducing the number of generation steps. By breaking away from conventional sequential modeling paradigms, NAR achieves 2.4× and 8.6× throughput improvements on ImageNet and UCF101, respectively, while attaining superior FID/FVD scores. A compact 0.8B-parameter model surpasses Chameleon-7B on GenEval and reduces required training data by 60% (to 40% of the baseline).
Masked autoregressive (MAR) models suffer from limited inference acceleration due to the need for sequential, single-step modeling of highly spatially correlated visual tokens. To address this, we propose a Generate-then-Reconstruct (GtR) two-stage sampling paradigm: first, a coarse-grained global structure is generated; second, a Frequency-guided Token Selection (FTS) mechanism—based on spectral energy analysis—selectively reconstructs high-frequency detail regions. GtR requires no additional training and integrates hierarchical sampling with off-the-shelf MAR models for multi-stage inference, effectively decoupling structural and textural modeling. Evaluated on ImageNet and text-to-image generation, GtR achieves a 3.72× speedup over standard MAR inference while attaining competitive fidelity (FID = 1.59) and diversity (Inception Score = 304.4), substantially outperforming existing acceleration methods without compromising generation quality.
Existing autoregressive vision generation models suffer from inefficiency in token-by-token decoding or complexity in multi-scale modeling. This paper proposes Expanding Autoregressive Representation (EAR), a novel framework inspired by human centric-outward visual perception, which employs a spiral token expansion order to explicitly model spatial continuity. EAR integrates parallel autoregressive decoding with a length-adaptive mechanism that dynamically adjusts the number of tokens predicted per step—thereby jointly optimizing generation quality, inference speed, and perceptual relevance. Experiments on ImageNet demonstrate that EAR achieves, for the first time within a single-scale autoregressive framework, a Pareto-optimal trade-off between synthesis fidelity and inference efficiency, significantly outperforming state-of-the-art methods. This work establishes a new paradigm for efficient, scalable, and cognitively aligned autoregressive visual modeling.
Autoregressive text-to-image models suffer from slow inference due to sequential token decoding—requiring thousands of forward passes. To address this, we propose a parallel denoising decoding framework in the embedding space: modeling denoising as Jacobi iteration, enabling simultaneous prediction of multiple clean tokens from a noise-initialized sequence; integrating a next-slice prediction paradigm and denoising-guided iterative trajectory. Our method requires only lightweight fine-tuning of pretrained models and incorporates probabilistic validation with progressive refinement to ensure generation stability. Experiments demonstrate substantial reduction in forward-pass count and significant speedup in inference, while preserving visual fidelity. The core contribution is the first principled integration of denoising principles with Jacobi-style parallel iteration, enabling efficient, stable, and transferable parallel decoding for autoregressive models.
To address the prohibitive computational and memory overhead in high-resolution autoregressive image generation—caused by quadratic growth of token count with resolution—this paper proposes an entropy-aware dynamic patching mechanism. It introduces the prediction entropy of a lightweight, unsupervised autoregressive model as an information-driven criterion to adaptively merge tokens into variable-sized image patches. The method supports dynamic patching during training and seamless inference-time scaling to larger patch sizes, while remaining fully compatible with standard decoder-only architectures. Key components include entropy-guided token aggregation, variable-length patch embedding, and multi-scale robust representation learning. On ImageNet, our approach reduces token counts by 1.81× and 2.06× at 256×256 and 384×384 resolutions, respectively, cuts training FLOPs by up to 40%, accelerates convergence, and improves FID by 27.1% over the baseline.
Visual autoregressive (AR) image generation faces a fundamental trade-off: random token ordering enables bidirectional context modeling but violates spatial priors (e.g., centrality bias and locality), whereas fixed raster-scan ordering respects such priors yet hinders flexible, region-specific editing. This work proposes Spanning Tree Autoregressive (STAR), a novel AR framework that constructs a structured token sequence via breadth-first traversal of a uniformly sampled spanning tree over the image patch lattice. STAR integrates spatial inductive biases intrinsically while preserving compatibility with standard language-style AR architectures. It supports efficient suffix completion via rejection sampling and enables dynamic specification—during inference—of both initial seed regions and edit masks. Experiments demonstrate that STAR significantly outperforms random permutation baselines on image editing tasks, achieving superior balance among sampling efficiency, contextual expressivity, and sequence-order flexibility without compromising generation quality.
Autoregressive image generation suffers from low information density and spatially uneven distribution of image tokens, limiting both generation quality and decoding speed. To address this, we propose an entropy-driven efficient decoding framework. Our key contributions are: (1) a spatial entropy-guided dynamic temperature scheduling mechanism that balances token diversity and structural consistency; (2) an entropy-aware acceptance criterion for speculative decoding, significantly improving token acceptance reliability; and (3) a lightweight design compatible with both mask-based and scale-wise autoregressive architectures. Extensive experiments across multiple benchmarks and model variants demonstrate that our method preserves near-lossless generation quality while reducing inference cost to 85% of conventional acceleration approaches—outperforming existing decoding strategies in both efficiency and fidelity.