autoregressive diffusion distillation

Design and implement distillation methods that convert autoregressive diffusion generative models into much faster samplers or autoregressive-style models suitable for real-time and streaming (frame- or chunk-wise) generation. This work includes crafting training objectives (e.g., combining teacher-forcing and self-forcing), algorithms to replace iterative diffusion sampling with efficient inference, and evaluating the fidelity/latency tradeoffs of the distilled models.

autoregressivediffusiondistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing video generation methods struggle to simultaneously achieve high fidelity, motion coherence, and low latency in streaming scenarios, with long-sequence generation particularly susceptible to error accumulation. This work proposes Diagonal Distillation, a novel approach that employs an asymmetric generation strategy—performing multi-step denoising in early stages and few-step synthesis in later stages—to explicitly align noise prediction with inference conditions, thereby mitigating exposure bias. The method further integrates implicit optical flow modeling and temporal context conditioning to enable efficient and coherent autoregressive video synthesis. Evaluated on 5-second video generation, the approach achieves a runtime of only 2.61 seconds (up to 31 FPS), yielding a 277.3× speedup over the original diffusion model while significantly improving visual quality and motion consistency in long sequences.

autoregressivediffusion distillationstreaming

Existing bidirectional video diffusion models suffer from prohibitive computational overhead due to global frame attention, hindering real-time deployment in interactive applications such as gaming. To address this, we propose the first causal autoregressive video diffusion Transformer, overcoming the error accumulation bottleneck inherent in autoregressive generation via three key innovations: (1) Distribution-Matching Distillation (DMD) in the video domain, enabling faithful trajectory-level knowledge transfer from a teacher to a student model; (2) ODE-trajectory-guided student initialization, improving the plausibility of initial latent states; and (3) an asymmetric causal teacher–student supervision scheme, balancing generation fidelity and inference efficiency. Our model achieves streaming video synthesis at 9.4 FPS on a single GPU with only four sampling steps—down from 50—while setting a new SOTA score of 84.27 on VBench-Long. It supports zero-shot long-video generation and multimodal streaming tasks including video-to-video and image-to-video translation.

Bidirectional Attention MechanismReal-time ProcessingVideo Production Models

This work addresses the limitations of existing autoregressive diffusion distillation methods, which suffer from coarse response granularity and high sampling latency, hindering real-time interactive video generation at an ultra-low latency of 1–2 steps per frame. To overcome this, the authors propose Causal Forcing++, introducing a novel causal consistency distillation mechanism that enables high-quality video synthesis in just 1–2 steps per frame within an autoregressive framework, effectively balancing low latency and controllability. The method eliminates the need for precomputing full ODE trajectories by integrating online teacher ODE single-step supervision with autoregressive conditional flow mapping, substantially improving training efficiency and stability. Experiments demonstrate that under a 2-step-per-frame setting, Causal Forcing++ achieves gains of 0.1, 0.3, and 0.335 in VBench Total, Quality, and VisionReward scores, respectively, reduces first-frame latency by 50%, and cuts second-stage training cost by approximately fourfold.

few-step autoregressive diffusionframe-wise autoregressionreal-time interactive video generation

This work addresses the inefficiency and suboptimal generation quality of autoregressive video diffusion models in streaming generation and action-conditioned interactive world modeling. To overcome these limitations, the authors propose a unified causal diffusion distillation framework that integrates teacher forcing and self-forcing strategies, and—critically—introduces continuous-time rectified consistency models (rCM) into autoregressive video generation for the first time. The resulting approach, termed Causal-rCM, leverages a causal diffusion Transformer, distribution matching distillation (DMD), and a custom masked FlashAttention-2 Jacobian-vector product kernel to enable scalable and efficient training. Remarkably, the 2-step sampling Wan2.1-1.3B model, trained solely on synthetic data, achieves a score of 84.63 on VBench-T2V, substantially outperforming existing methods, and has been successfully integrated into Cosmos 3 for high-efficiency interactive world modeling.

autoregressive video diffusioncausal trainingdiffusion distillation

Adaptive Non-Uniform Timestep Sampling for Diffusion Model Training

Nov 15, 2024
MK
Myunsoo Kim
🏛️ Korea University

In diffusion model training, uniform timestep sampling ignores the variance heterogeneity of gradients across timesteps, rendering high-variance timesteps convergence bottlenecks. To address this, we propose the first online evaluation mechanism that dynamically assesses—per iteration—the impact of gradient updates on the objective function, enabling adaptive, non-uniform timestep sampling focused on optimization-sensitive timesteps. Our method transcends conventional static weighting or heuristic sampling by unifying gradient variance analysis, objective impact tracking, and importance sampling. It achieves principled, real-time timestep prioritization without requiring precomputed statistics or architectural modifications. Experiments across diverse datasets, noise schedules, and network architectures demonstrate consistent improvements: our approach accelerates convergence and enhances final model performance compared to state-of-the-art timestep sampling and weighting strategies.

Accelerates convergence while improving final model performanceAddresses high gradient variance in diffusion model trainingIntroduces adaptive non-uniform timestep sampling method

Latest Papers

What's happening recently
View more

ReDiF: Reinforced Distillation for Few Step Diffusion

Dec 28, 2025
AT
Amirhossein Tighkhorshid
🏛️ Sharif University of Technology | Alan Turing Institute | London School of Economics

To address the high sampling cost and slow inference of diffusion models, this paper proposes the first framework that formulates knowledge distillation as a reinforcement learning (RL) policy optimization problem. Specifically, the student model’s multi-step denoising process is modeled as a sequential decision-making task, with sparse reward signals derived from alignment to teacher outputs; the Proximal Policy Optimization (PPO) algorithm is employed to optimize long-step denoising policies. Unlike conventional distillation paradigms—characterized by fixed step counts and layer-wise matching—our approach supports dynamic step scheduling and is model-agnostic, enabling general-purpose distillation. The reward design is inherently extensible. Experiments demonstrate that our method achieves superior generation quality over existing distillation approaches using only 4–8 inference steps, while exhibiting strong generalization and stability across multiple benchmark datasets.

Accelerates diffusion model sampling via reinforcement learning distillationOptimizes student models to take longer denoising steps efficientlyReduces inference steps and computational cost in diffusion models

This work addresses the performance degradation and lack of theoretical guarantees in existing approaches when distilling bidirectional video diffusion models into autoregressive ones, primarily due to architectural mismatches. The authors propose Causal Forcing, a novel method that leverages an autoregressive teacher model to guide the initialization of the ordinary differential equation (ODE) solver, effectively bridging the architectural gap between bidirectional and causal attention mechanisms. This approach provides the first theoretical resolution to the frame-level non-injectivity issue inherent in autoregressive distillation, ensuring invertibility of the flow map and circumventing performance loss caused by conditional expectation solutions. Built upon an ODE-based distillation framework and integrating autoregressive diffusion with causal attention, the method achieves state-of-the-art results, surpassing prior art by 19.3%, 8.7%, and 16.7% on Dynamic Degree, VisionReward, and Instruction Following metrics, respectively, enabling high-quality, real-time interactive video generation.

architectural gapautoregressive diffusioncausal attention

This work addresses the inefficiency of autoregressive video diffusion models, whose fixed denoising schedules often lead to either computational redundancy or insufficient refinement, hindering real-time generation. To overcome this limitation, the authors propose DSA, a confidence-guided dynamic computation framework that introduces adaptive step allocation into such models for the first time. DSA employs a lightweight confidence head to predict the reliability of denoising at each frame and jointly trains the generator and confidence head using a distribution-matching distillation objective. During inference, the model adaptively terminates or continues denoising based on predicted confidence, without requiring additional data or heuristic rules. Evaluated on an H100 GPU, the method achieves 22.63 FPS with sub-second latency while matching or surpassing state-of-the-art models in VBench quality metrics.

adaptive computationautoregressive video generationdiffusion models

Existing distillation methods for video diffusion models directly adopt image-based distillation strategies, often leading to issues such as oversaturation, temporal inconsistency, and mode collapse. This work proposes the first distillation framework specifically designed for video diffusion models, featuring an adaptive regression loss that dynamically modulates spatial supervision strength and a temporal regularization loss to suppress inter-frame inconsistencies. Coupled with a frame interpolation strategy during inference, the method enables efficient, high-quality video generation at extremely low sampling steps. Evaluated on the VBench and VBench2 benchmarks, the proposed approach significantly outperforms existing distillation techniques, achieving superior perceptual fidelity, natural motion dynamics, and stable few-step synthesis.

few-step generationoversaturationtemporal collapse

Hot Scholars

LS

Lingyun Sun

Zhejiang University
Design IntelligenceHCIArtificial IntelligenceIndustrial Design
YZ

Yifan Zhan

The University of Tokyo
3D Vision
ZW

Zhixiang Wang

University of Tokyo
Computational PhotographyComputational ImagingMachine Learning
ZX

Zeke Xie

Assistant Professor, The Hong Kong University of Science and Technology (Guangzhou)/ PI, xLeaF Lab
Generative AIData-centric AILarge ModelsDeep Learning Theory