visual autoregressive denoising

Designs, builds, or evaluates autoregressive generative models that reconstruct or denoise visual data by sequentially predicting pixels or discrete visual tokens; these systems produce progressive, prefix-conditioned or multiscale token maps and model global-to-local structural dependencies via next-token/next-scale probability predictions.

visualautoregressivedenoising

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.54
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the longstanding trade-off between generation quality and reconstruction fidelity in autoregressive image synthesis, which traditionally relies on visual tokenizers trained in a separate, multi-stage pipeline. To overcome this limitation, the authors propose the first end-to-end joint training framework that simultaneously optimizes a 1D semantic tokenizer and an autoregressive generative model. Central to their approach is the integration of a vision foundation model to enrich the semantic expressiveness of the learned tokens, along with a novel mechanism that directly supervises tokenizer learning through the generated outputs—thereby departing from conventional two-stage paradigms. Evaluated on class-unconditional ImageNet generation at 256×256 resolution, the method achieves a new state-of-the-art FID of 1.48.

1D semantic tokenizerautoregressive image generationend-to-end training

Next Patch Prediction for Autoregressive Visual Generation

Dec 19, 2024
YP
Yatian Pang
🏛️ Peking University | NUS | UCF | HKUST | PengCheng Laboratory

To address the challenges of excessively long sequences, high training costs, and difficulty in leveraging intrinsic image hierarchies in autoregressive image generation, this paper proposes the Next Patch Prediction (NPP) paradigm. NPP aggregates low-level image tokens into high-information-density patch tokens, drastically reducing sequence length; it introduces a novel multi-scale, coarse-to-fine hierarchical patch grouping strategy that requires no model architecture modifications, additional parameters, or custom tokenizers—ensuring generality and plug-and-play compatibility. Evaluated under standard autoregressive modeling and FID assessment on ImageNet, NPP reduces training cost to approximately 0.6× that of baseline methods while improving FID by up to 1.0. The improvement is consistent across models ranging from 100M to 1.4B parameters. The core contribution is the first seamless integration of inherent image hierarchy into autoregressive generation—achieving simultaneous gains in efficiency, sample quality, and deployment practicality.

Hierarchical ImprovementsHigh-quality Image GenerationTraining Efficiency

Parallelized Autoregressive Visual Generation

Dec 19, 2024
YW
Yuqing Wang
🏛️ University of Hong Kong | Peking University | ByteDance

Autoregressive visual generation suffers from low inference efficiency due to sequential, token-by-token decoding. This paper proposes a dependency-aware parallelization strategy that requires no modification to the model architecture or tokenizer. It dynamically partitions the decoding process into parallel and serial regions based on the conditional dependency strength among visual tokens—marking the first integration of explicit dependency modeling into parallel decoding decisions. The method comprises dependency-aware grouping sampling and localized serial constraints, and is plug-and-play on standard Transformer decoders. Evaluated on image and video generation tasks using ImageNet and UCF-101, it achieves up to 9.5× speedup; critically, it attains 3.6× acceleration without compromising generation quality—outperforming existing parallel decoding approaches. Its core contributions are: (i) zero-modification deployment, (ii) dependency-driven parallelization, and (iii) a general, efficient, and quality-preserving acceleration paradigm for autoregressive visual generation.

Balancing parallel and sequential token generation efficientlyMaintaining quality while accelerating image and video generationSlow inference speed in autoregressive visual generation models

Neighboring Autoregressive Modeling for Efficient Visual Generation

Mar 12, 2025
YH
Yefei He
🏛️ Zhejiang University | Shanghai AI Laboratory | The University of Adelaide

Existing vision autoregressive models rely on raster-scan ordering for “next-token prediction,” neglecting the intrinsic spatial-temporal locality of visual data. This work proposes Neighborhood Autoregressive Modeling (NAR), reformulating generation as a progressive outpainting process ordered by increasing Manhattan distance from a seed region. It introduces a novel “next-neighborhood prediction” mechanism and a multi-dimensional orthogonal decoding head, enabling parallel prediction of spatial-temporal neighborhood tokens and drastically reducing the number of generation steps. By breaking away from conventional sequential modeling paradigms, NAR achieves 2.4× and 8.6× throughput improvements on ImageNet and UCF101, respectively, while attaining superior FID/FVD scores. A compact 0.8B-parameter model surpasses Chameleon-7B on GenEval and reduces required training data by 60% (to 40% of the baseline).

Achieves higher throughput and better quality in image and video generation.Improves visual generation efficiency by focusing on spatial-temporal locality.Introduces parallel prediction of adjacent tokens to reduce generation steps.

Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage Sampling

Oct 20, 2025
FY
Feihong Yan
🏛️ SJTU | Tsinghua University | Beijing Institute of Technology

Masked autoregressive (MAR) models suffer from limited inference acceleration due to the need for sequential, single-step modeling of highly spatially correlated visual tokens. To address this, we propose a Generate-then-Reconstruct (GtR) two-stage sampling paradigm: first, a coarse-grained global structure is generated; second, a Frequency-guided Token Selection (FTS) mechanism—based on spectral energy analysis—selectively reconstructs high-frequency detail regions. GtR requires no additional training and integrates hierarchical sampling with off-the-shelf MAR models for multi-stage inference, effectively decoupling structural and textural modeling. Evaluated on ImageNet and text-to-image generation, GtR achieves a 3.72× speedup over standard MAR inference while attaining competitive fidelity (FID = 1.59) and diversity (Inception Score = 304.4), substantially outperforming existing acceleration methods without compromising generation quality.

Accelerating masked autoregressive models via two-stage hierarchical samplingOptimizing token selection based on frequency-weighted semantic importanceReducing computational complexity while maintaining image generation quality

Latest Papers

What's happening recently
View more

Learning to Expand Images for Efficient Visual Autoregressive Modeling

Nov 19, 2025
RY
Ruiqing Yang
🏛️ University of Electronic Science and Technology of China | Central South University | Xidian University | SenseTime Research | Shanghai Jiao Tong University

Existing autoregressive vision generation models suffer from inefficiency in token-by-token decoding or complexity in multi-scale modeling. This paper proposes Expanding Autoregressive Representation (EAR), a novel framework inspired by human centric-outward visual perception, which employs a spiral token expansion order to explicitly model spatial continuity. EAR integrates parallel autoregressive decoding with a length-adaptive mechanism that dynamically adjusts the number of tokens predicted per step—thereby jointly optimizing generation quality, inference speed, and perceptual relevance. Experiments on ImageNet demonstrate that EAR achieves, for the first time within a single-scale autoregressive framework, a Pareto-optimal trade-off between synthesis fidelity and inference efficiency, significantly outperforming state-of-the-art methods. This work establishes a new paradigm for efficient, scalable, and cognitively aligned autoregressive visual modeling.

Complexity issues with multi-scale representations in image generationInefficient token-by-token decoding in visual autoregressive modelsPoor alignment between generation order and perceptual relevance patterns

Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation

Oct 10, 2025
YT
Yao Teng
🏛️ The University of Hong Kong | CUHK | Huawei Noah's Ark Lab | Tsinghua University

Autoregressive text-to-image models suffer from slow inference due to sequential token decoding—requiring thousands of forward passes. To address this, we propose a parallel denoising decoding framework in the embedding space: modeling denoising as Jacobi iteration, enabling simultaneous prediction of multiple clean tokens from a noise-initialized sequence; integrating a next-slice prediction paradigm and denoising-guided iterative trajectory. Our method requires only lightweight fine-tuning of pretrained models and incorporates probabilistic validation with progressive refinement to ensure generation stability. Experiments demonstrate substantial reduction in forward-pass count and significant speedup in inference, while preserving visual fidelity. The core contribution is the first principled integration of denoising principles with Jacobi-style parallel iteration, enabling efficient, stable, and transferable parallel decoding for autoregressive models.

Accelerates slow autoregressive image generationEnables parallel token prediction via denoisingReduces model forward passes while maintaining quality

DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation

Dec 26, 2025
DS
Divyansh Srivastava
🏛️ University of California, San Diego | Dolby Laboratories

To address the prohibitive computational and memory overhead in high-resolution autoregressive image generation—caused by quadratic growth of token count with resolution—this paper proposes an entropy-aware dynamic patching mechanism. It introduces the prediction entropy of a lightweight, unsupervised autoregressive model as an information-driven criterion to adaptively merge tokens into variable-sized image patches. The method supports dynamic patching during training and seamless inference-time scaling to larger patch sizes, while remaining fully compatible with standard decoder-only architectures. Key components include entropy-guided token aggregation, variable-length patch embedding, and multi-scale robust representation learning. On ImageNet, our approach reduces token counts by 1.81× and 2.06× at 256×256 and 384×384 resolutions, respectively, cuts training FLOPs by up to 40%, accelerates convergence, and improves FID by 27.1% over the baseline.

Dynamically aggregates tokens into variable-sized patches for efficiencyImproves training efficiency and image quality with adaptive token mergingReduces computational and memory demands in autoregressive image generation

Spanning Tree Autoregressive Visual Generation

Nov 21, 2025
SL
Sangkyu Lee
🏛️ Yonsei University | LG AI Research | KIST | University of Michigan, Ann Arbor | Seoul National University

Visual autoregressive (AR) image generation faces a fundamental trade-off: random token ordering enables bidirectional context modeling but violates spatial priors (e.g., centrality bias and locality), whereas fixed raster-scan ordering respects such priors yet hinders flexible, region-specific editing. This work proposes Spanning Tree Autoregressive (STAR), a novel AR framework that constructs a structured token sequence via breadth-first traversal of a uniformly sampled spanning tree over the image patch lattice. STAR integrates spatial inductive biases intrinsically while preserving compatibility with standard language-style AR architectures. It supports efficient suffix completion via rejection sampling and enables dynamic specification—during inference—of both initial seed regions and edit masks. Experiments demonstrate that STAR significantly outperforms random permutation baselines on image editing tasks, achieving superior balance among sampling efficiency, contextual expressivity, and sequence-order flexibility without compromising generation quality.

Incorporating image priors like center bias and locality into autoregressive visual generationMaintaining sampling performance while enabling flexible sequence orders for image editingOvercoming performance decline in bidirectional context modeling with structured randomized strategies

Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy

Oct 10, 2025
XM
Xiaoxiao Ma
🏛️ University of Science and Technology of China | Meituan

Autoregressive image generation suffers from low information density and spatially uneven distribution of image tokens, limiting both generation quality and decoding speed. To address this, we propose an entropy-driven efficient decoding framework. Our key contributions are: (1) a spatial entropy-guided dynamic temperature scheduling mechanism that balances token diversity and structural consistency; (2) an entropy-aware acceptance criterion for speculative decoding, significantly improving token acceptance reliability; and (3) a lightweight design compatible with both mask-based and scale-wise autoregressive architectures. Extensive experiments across multiple benchmarks and model variants demonstrate that our method preserves near-lossless generation quality while reducing inference cost to 85% of conventional acceleration approaches—outperforming existing decoding strategies in both efficiency and fidelity.

Addressing low information density in image tokensEnhancing sampling efficiency with entropy-guided decodingImproving autoregressive image generation quality and speed

Hot Scholars

BW

Bihan Wen

Associate Professor, Nanyang Technological University
Machine LearningImage ProcessingComputational ImagingComputer Vision
GP

Gjergj Plepi

PhD student, Autonomous Intelligent Systems, University of Bonn
Computer VisionRepresentation LearningRobotics
SP

Sihwan Park

KAIST, Ph.D student
Deep LearningMachine LearningOptimizationLarge Language Models
SZ

Shaoting Zhang

Shanghai AI Lab; SenseTime Research
Medical Image AnalysisComputer VisionFoundation Models
DJ

Doohyuk Jang

KAIST AI
Machine LearningNatural Language Processing