Score
Designs, implements, and evaluates knowledge-distillation procedures that transfer behavior from high-capacity autoregressive image generative models into compact or faster student autoregressive models. This work includes methods that train on student-generated samples, selectively apply teacher supervision, modify token-level losses to reduce prediction ambiguity, and measure or optimize long-horizon generation quality.
This work addresses the limited efficacy of existing knowledge distillation methods for visual autoregressive (AR) image generation models, which stems from long-sequence decoding and ambiguity in visual tokens. The study presents the first systematic investigation into distillation mechanisms for visual AR models and introduces VarKD, a novel framework that selects high-quality samples via a student-aware sampling strategy and applies selective teacher supervision at the token level to mitigate unreliable supervisory signals. Tailored to the characteristics of image generation, VarKD consistently outperforms current distillation approaches across diverse AR backbone architectures on ImageNet, substantially narrowing the performance gap between compact student models and large-scale teacher models.
This work addresses the longstanding trade-off between generation quality and reconstruction fidelity in autoregressive image synthesis, which traditionally relies on visual tokenizers trained in a separate, multi-stage pipeline. To overcome this limitation, the authors propose the first end-to-end joint training framework that simultaneously optimizes a 1D semantic tokenizer and an autoregressive generative model. Central to their approach is the integration of a vision foundation model to enrich the semantic expressiveness of the learned tokens, along with a novel mechanism that directly supervises tokenizer learning through the generated outputs—thereby departing from conventional two-stage paradigms. Evaluated on class-unconditional ImageNet generation at 256×256 resolution, the method achieves a new state-of-the-art FID of 1.48.
Autoregressive image generation models, which adapt the NLP paradigm of “next-token prediction” to vision, face three fundamental bottlenecks: insufficient modeling of local and conditional dependencies, inter-step semantic inconsistency, and lack of spatial invariance—hindering high-level visual semantic learning. To address these, we propose a “comprehend-then-generate” self-supervised training paradigm, introducing Self-guided Training for AutoRegressive models (ST-AR)—a framework requiring no external pretrained models. Its core innovation lies in a novel self-supervised objective that explicitly guides the model to learn structured semantic representations *before* autoregressive decoding. Evaluated on the LlamaGen family, ST-AR achieves up to 42% (LlamaGen-L) and 49% (LlamaGen-XL) FID improvement, significantly enhancing generation quality while preserving full compatibility with existing sampling strategies. This work is the first to systematically identify and resolve the intrinsic limitations of next-token prediction in visual generative modeling.
This work addresses the challenge of efficiently generating high-quality training data required for supervised learning in text-to-image generation models. We propose the Guided Adversarial Prompts (GAP) framework—a closed-loop data generation system integrating three core mechanisms: (1) adversarial prompt optimization guided by supervised model loss, (2) target distribution alignment via feature matching or discriminator-based guidance, and (3) online feedback adaptation. GAP is the first method to synergistically couple adversarial generation with explicit distributional constraints, shifting data synthesis from open-loop, static prompting to closed-loop, adaptive refinement. Empirical evaluation across diverse settings—including multi-task learning, heterogeneous model architectures, and distribution shifts (e.g., spurious correlations, unseen domains)—demonstrates substantial improvements in downstream model generalization. Data utilization efficiency increases by up to 3.2× compared to baseline approaches.
In knowledge distillation, supervised methods suffer from train-inference distribution mismatch, while on-policy approaches yield inaccurate teacher feedback due to low-quality student-generated samples. This paper proposes Speculative Distillation—a novel framework where the student first generates candidate token sequences, and the teacher dynamically corrects only low-confidence tokens, enabling high-fidelity knowledge transfer under inference-time distribution alignment. Its core innovation is the first online, token-level, teacher-student collaborative correction mechanism, integrating confidence-driven interleaved sampling, teacher-guided dynamic reweighting, and multi-task joint training. Evaluated across machine translation, summarization, mathematical reasoning, and instruction-following tasks, the method consistently outperforms both supervised and on-policy distillation baselines. It demonstrates robust performance gains across diverse model scales, data regimes, and initialization strategies.
Existing video distillation methods rely on non-causal teacher models whose supervision incorporates future frames and control signals, creating a mismatch with the causal, autoregressive generation process of student models and leading to a train-inference discrepancy. This work proposes Context-Matched Distillation (CMD), a novel framework that introduces strict causal alignment in video distillation for the first time. CMD employs a causal teacher model that scores target frames using only past and current information, and integrates two key strategies: Prefix Scoring, which aligns the teacher’s context with the actual generation context of the student, and Prefix Corruption, which stabilizes early-stage training. This approach ensures consistent causal modeling between teacher and student during both training and inference. CMD supports frame-level, block-level, long-form video generation, and time-varying camera-conditioned synthesis, achieving state-of-the-art performance in autoregressive video generation and significantly improving adherence to dynamic camera controls across multiple benchmarks.
This work addresses the long-standing absence of a unified statistical perspective on knowledge distillation, which has frequently been perceived as an engineering heuristic. We propose a unifying framework grounded in Bayesian inference that formalizes teacher model predictions as prior information, thereby enabling principled uncertainty quantification. This framework not only bridges classical distillation methods with their extensions to large language models but also integrates seamlessly with modern generative systems. Furthermore, we provide a conceptual roadmap and identify key open problems, establishing a systematic foundation for deepening the theoretical understanding of distillation mechanisms.
This work addresses the degradation in classification performance caused by insufficiently discriminative patterns in short temporal inputs and the difficulty of effectively transferring knowledge from long-context teacher models. To this end, it introduces diffusion priors into the knowledge distillation framework for the first time. The method treats the student model’s short-context features as degraded observations of the teacher’s long-context representations and leverages a diffusion model to generate diverse, long-context supervisory signals. Task-relevant knowledge is then transferred through Bayesian posterior sampling, enabling distributed and adaptive distillation. Extensive experiments demonstrate that the proposed approach significantly improves short-sequence classification accuracy across various early-exit configurations, datasets, and model architectures, effectively narrowing the generalization gap induced by input length discrepancies.
This work addresses the prevailing focus on distillation objectives in existing few-step distillation methods, which often overlooks the critical influence of the training pipeline on student model performance. Taking Qwen-Image-2.0 as the baseline, the study systematically investigates the interplay among three key training components—data composition, teacher guidance, and task mixing—in both text-to-image generation and instruction-guided image editing tasks, thereby transcending the limitations of solely optimizing distillation targets. The authors propose a more holistic few-step distillation paradigm by integrating multi-task mixed training, refined data ratio scheduling, and dynamic teacher guidance strategies. Experimental results demonstrate that this approach substantially enhances student model performance across both generation and editing tasks, underscoring the decisive role of training pipeline design in effective knowledge distillation.
Training large-scale vision models is computationally expensive, and existing knowledge distillation methods primarily focus on model compression or accuracy improvement rather than accelerating the training of strong models. This work proposes a plug-and-play weak-to-strong knowledge distillation strategy that leverages a fixed-weight weak teacher model during early training stages and dynamically terminates distillation once the student surpasses the teacher’s performance, significantly reducing the number of epochs needed to reach target accuracy. Notably, this is the first approach to employ knowledge distillation explicitly for accelerating strong model training rather than compression. The method demonstrates broad applicability across image classification, object detection, and diffusion-based generation tasks, achieving up to 4.8× epoch acceleration on ImageNet and CIFAR, 1.7× speedup on COCO detection, and a 2.5× reduction in FID-convergent steps for CIFAR-10 diffusion models.