autoregressive model distillation

Designs, implements, and evaluates knowledge-distillation procedures that transfer behavior from high-capacity autoregressive image generative models into compact or faster student autoregressive models. This work includes methods that train on student-generated samples, selectively apply teacher supervision, modify token-level losses to reduce prediction ambiguity, and measure or optimize long-horizon generation quality.

autoregressivemodeldistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$207K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limited efficacy of existing knowledge distillation methods for visual autoregressive (AR) image generation models, which stems from long-sequence decoding and ambiguity in visual tokens. The study presents the first systematic investigation into distillation mechanisms for visual AR models and introduces VarKD, a novel framework that selects high-quality samples via a student-aware sampling strategy and applies selective teacher supervision at the token level to mitigate unreliable supervisory signals. Tailored to the characteristics of image generation, VarKD consistently outperforms current distillation approaches across diverse AR backbone architectures on ImageNet, substantially narrowing the performance gap between compact student models and large-scale teacher models.

Image GenerationKnowledge DistillationModel Compression

This work addresses the longstanding trade-off between generation quality and reconstruction fidelity in autoregressive image synthesis, which traditionally relies on visual tokenizers trained in a separate, multi-stage pipeline. To overcome this limitation, the authors propose the first end-to-end joint training framework that simultaneously optimizes a 1D semantic tokenizer and an autoregressive generative model. Central to their approach is the integration of a vision foundation model to enrich the semantic expressiveness of the learned tokens, along with a novel mechanism that directly supervises tokenizer learning through the generated outputs—thereby departing from conventional two-stage paradigms. Evaluated on class-unconditional ImageNet generation at 256×256 resolution, the method achieves a new state-of-the-art FID of 1.48.

1D semantic tokenizerautoregressive image generationend-to-end training

Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation

Sep 18, 2025
XY
Xiaoyu Yue
🏛️ Shanghai AI Laboratory | Chinese University of Hong Kong | University of Hong Kong | University of Sydney

Autoregressive image generation models, which adapt the NLP paradigm of “next-token prediction” to vision, face three fundamental bottlenecks: insufficient modeling of local and conditional dependencies, inter-step semantic inconsistency, and lack of spatial invariance—hindering high-level visual semantic learning. To address these, we propose a “comprehend-then-generate” self-supervised training paradigm, introducing Self-guided Training for AutoRegressive models (ST-AR)—a framework requiring no external pretrained models. Its core innovation lies in a novel self-supervised objective that explicitly guides the model to learn structured semantic representations *before* autoregressive decoding. Evaluated on the LlamaGen family, ST-AR achieves up to 42% (LlamaGen-L) and 49% (LlamaGen-XL) FID improvement, significantly enhancing generation quality while preserving full compatibility with existing sampling strategies. This work is the first to systematically identify and resolve the intrinsic limitations of next-token prediction in visual generative modeling.

Addressing autoregressive models' visual semantics learning limitationsEnhancing generation quality via self-supervised training objectivesImproving image understanding without pre-trained representation models

Controlled Training Data Generation with Diffusion Models

Mar 22, 2024
TY
Teresa Yeo
🏛️ Swiss Federal Institute of Technology Lausanne | MIT

This work addresses the challenge of efficiently generating high-quality training data required for supervised learning in text-to-image generation models. We propose the Guided Adversarial Prompts (GAP) framework—a closed-loop data generation system integrating three core mechanisms: (1) adversarial prompt optimization guided by supervised model loss, (2) target distribution alignment via feature matching or discriminator-based guidance, and (3) online feedback adaptation. GAP is the first method to synergistically couple adversarial generation with explicit distributional constraints, shifting data synthesis from open-loop, static prompting to closed-loop, adaptive refinement. Empirical evaluation across diverse settings—including multi-task learning, heterogeneous model architectures, and distribution shifts (e.g., spurious correlations, unseen domains)—demonstrates substantial improvements in downstream model generalization. Data utilization efficiency increases by up to 3.2× compared to baseline approaches.

Automate closed-loop feedback for adversarial prompt generationControl text-to-image models for supervised training dataGuide generation to match target data distributions

Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling

Oct 15, 2024
WX
Wenda Xu
🏛️ UC Santa Barbara | Google Cloud AI Research | CMU | Google DeepMind

In knowledge distillation, supervised methods suffer from train-inference distribution mismatch, while on-policy approaches yield inaccurate teacher feedback due to low-quality student-generated samples. This paper proposes Speculative Distillation—a novel framework where the student first generates candidate token sequences, and the teacher dynamically corrects only low-confidence tokens, enabling high-fidelity knowledge transfer under inference-time distribution alignment. Its core innovation is the first online, token-level, teacher-student collaborative correction mechanism, integrating confidence-driven interleaved sampling, teacher-guided dynamic reweighting, and multi-task joint training. Evaluated across machine translation, summarization, mathematical reasoning, and instruction-following tasks, the method consistently outperforms both supervised and on-policy distillation baselines. It demonstrates robust performance gains across diverse model scales, data regimes, and initialization strategies.

Addresses teacher-student knowledge gap in distillation.Enhances student model performance across diverse tasks.Improves training data quality in knowledge distillation.

Latest Papers

What's happening recently
View more

Existing video distillation methods rely on non-causal teacher models whose supervision incorporates future frames and control signals, creating a mismatch with the causal, autoregressive generation process of student models and leading to a train-inference discrepancy. This work proposes Context-Matched Distillation (CMD), a novel framework that introduces strict causal alignment in video distillation for the first time. CMD employs a causal teacher model that scores target frames using only past and current information, and integrates two key strategies: Prefix Scoring, which aligns the teacher’s context with the actual generation context of the student, and Prefix Corruption, which stabilizes early-stage training. This approach ensures consistent causal modeling between teacher and student during both training and inference. CMD supports frame-level, block-level, long-form video generation, and time-varying camera-conditioned synthesis, achieving state-of-the-art performance in autoregressive video generation and significantly improving adherence to dynamic camera controls across multiple benchmarks.

autoregressive video generationcausal distillationonline control

This work addresses the long-standing absence of a unified statistical perspective on knowledge distillation, which has frequently been perceived as an engineering heuristic. We propose a unifying framework grounded in Bayesian inference that formalizes teacher model predictions as prior information, thereby enabling principled uncertainty quantification. This framework not only bridges classical distillation methods with their extensions to large language models but also integrates seamlessly with modern generative systems. Furthermore, we provide a conceptual roadmap and identify key open problems, establishing a systematic foundation for deepening the theoretical understanding of distillation mechanisms.

Bayesian FormulationKnowledge DistillationLarge Language Models

This work addresses the degradation in classification performance caused by insufficiently discriminative patterns in short temporal inputs and the difficulty of effectively transferring knowledge from long-context teacher models. To this end, it introduces diffusion priors into the knowledge distillation framework for the first time. The method treats the student model’s short-context features as degraded observations of the teacher’s long-context representations and leverages a diffusion model to generate diverse, long-context supervisory signals. Task-relevant knowledge is then transferred through Bayesian posterior sampling, enabling distributed and adaptive distillation. Extensive experiments demonstrate that the proposed approach significantly improves short-sequence classification accuracy across various early-exit configurations, datasets, and model architectures, effectively narrowing the generalization gap induced by input length discrepancies.

generalization gapknowledge distillationlong-context transfer

This work addresses the prevailing focus on distillation objectives in existing few-step distillation methods, which often overlooks the critical influence of the training pipeline on student model performance. Taking Qwen-Image-2.0 as the baseline, the study systematically investigates the interplay among three key training components—data composition, teacher guidance, and task mixing—in both text-to-image generation and instruction-guided image editing tasks, thereby transcending the limitations of solely optimizing distillation targets. The authors propose a more holistic few-step distillation paradigm by integrating multi-task mixed training, refined data ratio scheduling, and dynamic teacher guidance strategies. Experimental results demonstrate that this approach substantially enhances student model performance across both generation and editing tasks, underscoring the decisive role of training pipeline design in effective knowledge distillation.

few-step distillationinstruction-guided image editingtext-to-image generation

Training large-scale vision models is computationally expensive, and existing knowledge distillation methods primarily focus on model compression or accuracy improvement rather than accelerating the training of strong models. This work proposes a plug-and-play weak-to-strong knowledge distillation strategy that leverages a fixed-weight weak teacher model during early training stages and dynamically terminates distillation once the student surpasses the teacher’s performance, significantly reducing the number of epochs needed to reach target accuracy. Notably, this is the first approach to employ knowledge distillation explicitly for accelerating strong model training rather than compression. The method demonstrates broad applicability across image classification, object detection, and diffusion-based generation tasks, achieving up to 4.8× epoch acceleration on ImageNet and CIFAR, 1.7× speedup on COCO detection, and a 2.5× reduction in FID-convergent steps for CIFAR-10 diffusion models.

knowledge distillationstrong studenttraining acceleration