Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage Sampling

📅 2025-10-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Masked autoregressive (MAR) models suffer from limited inference acceleration due to the need for sequential, single-step modeling of highly spatially correlated visual tokens. To address this, we propose a Generate-then-Reconstruct (GtR) two-stage sampling paradigm: first, a coarse-grained global structure is generated; second, a Frequency-guided Token Selection (FTS) mechanism—based on spectral energy analysis—selectively reconstructs high-frequency detail regions. GtR requires no additional training and integrates hierarchical sampling with off-the-shelf MAR models for multi-stage inference, effectively decoupling structural and textural modeling. Evaluated on ImageNet and text-to-image generation, GtR achieves a 3.72× speedup over standard MAR inference while attaining competitive fidelity (FID = 1.59) and diversity (Inception Score = 304.4), substantially outperforming existing acceleration methods without compromising generation quality.

Technology Category

Computer Vision: Generative Adversarial Networks (GANs) for VisionNatural Language Processing: GenerationMachine Learning: Deep Generative Models & Autoencoders

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Masked Autoregressive (MAR) models promise better efficiency in visual generation than autoregressive (AR) models for the ability of parallel generation, yet their acceleration potential remains constrained by the modeling complexity of spatially correlated visual tokens in a single step. To address this limitation, we introduce Generation then Reconstruction (GtR), a training-free hierarchical sampling strategy that decomposes generation into two stages: structure generation establishing global semantic scaffolding, followed by detail reconstruction efficiently completing remaining tokens. Assuming that it is more difficult to create an image from scratch than to complement images based on a basic image framework, GtR is designed to achieve acceleration by computing the reconstruction stage quickly while maintaining the generation quality by computing the generation stage slowly. Moreover, observing that tokens on the details of an image often carry more semantic information than tokens in the salient regions, we further propose Frequency-Weighted Token Selection (FTS) to offer more computation budget to tokens on image details, which are localized based on the energy of high frequency information. Extensive experiments on ImageNet class-conditional and text-to-image generation demonstrate 3.72x speedup on MAR-H while maintaining comparable quality (e.g., FID: 1.59, IS: 304.4 vs. original 1.59, 299.1), substantially outperforming existing acceleration methods across various model scales and generation tasks. Our codes will be released in https://github.com/feihongyan1/GtR.
Problem

Research questions and friction points this paper is trying to address.

Accelerating masked autoregressive models via two-stage hierarchical sampling
Reducing computational complexity while maintaining image generation quality
Optimizing token selection based on frequency-weighted semantic importance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Two-stage hierarchical sampling strategy for acceleration
Frequency-weighted token selection prioritizing image details
Training-free method maintaining quality while speeding generation
🔎 Similar Papers
2024-03-04Computer Vision and Pattern RecognitionCitations: 3