Score
Design and implement distillation methods that convert autoregressive diffusion generative models into much faster samplers or autoregressive-style models suitable for real-time and streaming (frame- or chunk-wise) generation. This work includes crafting training objectives (e.g., combining teacher-forcing and self-forcing), algorithms to replace iterative diffusion sampling with efficient inference, and evaluating the fidelity/latency tradeoffs of the distilled models.
Diffusion models (DMs) achieve state-of-the-art performance in text-to-image generation, yet their high memory footprint and computational cost—stemming from iterative sampling—hinder edge deployment. Knowledge distillation of pre-trained DMs has emerged as a key efficiency-enhancement strategy, but existing works lack systematic organization. This paper presents the first methodology-driven, structured survey of DM distillation, categorizing approaches into three paradigms: output-loss distillation, trajectory distillation, and adversarial distillation. We unify their formulations by integrating techniques from knowledge distillation, denoising trajectory fitting, adversarial training, and multi-step sampling approximation, clarifying underlying principles, boundaries, interconnections, and application scopes. Our analysis identifies fundamental challenges—including the sampling-fidelity trade-off and cross-paradigm integration—and proposes future directions such as scalable distillation frameworks. This work fills a critical gap by providing the first comprehensive, taxonomy-based review of DM distillation.
Existing video generation methods struggle to simultaneously achieve high fidelity, motion coherence, and low latency in streaming scenarios, with long-sequence generation particularly susceptible to error accumulation. This work proposes Diagonal Distillation, a novel approach that employs an asymmetric generation strategy—performing multi-step denoising in early stages and few-step synthesis in later stages—to explicitly align noise prediction with inference conditions, thereby mitigating exposure bias. The method further integrates implicit optical flow modeling and temporal context conditioning to enable efficient and coherent autoregressive video synthesis. Evaluated on 5-second video generation, the approach achieves a runtime of only 2.61 seconds (up to 31 FPS), yielding a 277.3× speedup over the original diffusion model while significantly improving visual quality and motion consistency in long sequences.
Existing bidirectional video diffusion models suffer from prohibitive computational overhead due to global frame attention, hindering real-time deployment in interactive applications such as gaming. To address this, we propose the first causal autoregressive video diffusion Transformer, overcoming the error accumulation bottleneck inherent in autoregressive generation via three key innovations: (1) Distribution-Matching Distillation (DMD) in the video domain, enabling faithful trajectory-level knowledge transfer from a teacher to a student model; (2) ODE-trajectory-guided student initialization, improving the plausibility of initial latent states; and (3) an asymmetric causal teacher–student supervision scheme, balancing generation fidelity and inference efficiency. Our model achieves streaming video synthesis at 9.4 FPS on a single GPU with only four sampling steps—down from 50—while setting a new SOTA score of 84.27 on VBench-Long. It supports zero-shot long-video generation and multimodal streaming tasks including video-to-video and image-to-video translation.
This work addresses the limitations of existing autoregressive diffusion distillation methods, which suffer from coarse response granularity and high sampling latency, hindering real-time interactive video generation at an ultra-low latency of 1–2 steps per frame. To overcome this, the authors propose Causal Forcing++, introducing a novel causal consistency distillation mechanism that enables high-quality video synthesis in just 1–2 steps per frame within an autoregressive framework, effectively balancing low latency and controllability. The method eliminates the need for precomputing full ODE trajectories by integrating online teacher ODE single-step supervision with autoregressive conditional flow mapping, substantially improving training efficiency and stability. Experiments demonstrate that under a 2-step-per-frame setting, Causal Forcing++ achieves gains of 0.1, 0.3, and 0.335 in VBench Total, Quality, and VisionReward scores, respectively, reduces first-frame latency by 50%, and cuts second-stage training cost by approximately fourfold.
This work addresses the inefficiency and suboptimal generation quality of autoregressive video diffusion models in streaming generation and action-conditioned interactive world modeling. To overcome these limitations, the authors propose a unified causal diffusion distillation framework that integrates teacher forcing and self-forcing strategies, and—critically—introduces continuous-time rectified consistency models (rCM) into autoregressive video generation for the first time. The resulting approach, termed Causal-rCM, leverages a causal diffusion Transformer, distribution matching distillation (DMD), and a custom masked FlashAttention-2 Jacobian-vector product kernel to enable scalable and efficient training. Remarkably, the 2-step sampling Wan2.1-1.3B model, trained solely on synthetic data, achieves a score of 84.63 on VBench-T2V, substantially outperforming existing methods, and has been successfully integrated into Cosmos 3 for high-efficiency interactive world modeling.
In diffusion model training, uniform timestep sampling ignores the variance heterogeneity of gradients across timesteps, rendering high-variance timesteps convergence bottlenecks. To address this, we propose the first online evaluation mechanism that dynamically assesses—per iteration—the impact of gradient updates on the objective function, enabling adaptive, non-uniform timestep sampling focused on optimization-sensitive timesteps. Our method transcends conventional static weighting or heuristic sampling by unifying gradient variance analysis, objective impact tracking, and importance sampling. It achieves principled, real-time timestep prioritization without requiring precomputed statistics or architectural modifications. Experiments across diverse datasets, noise schedules, and network architectures demonstrate consistent improvements: our approach accelerates convergence and enhances final model performance compared to state-of-the-art timestep sampling and weighting strategies.
本文提出了一种解耦自强迫蒸馏方法,用于流式说话头生成,通过在低维运动空间中融合条件并使用因果自回归变换器生成运动潜变量,提高了生成视频的质量和效率。
To address the high sampling cost and slow inference of diffusion models, this paper proposes the first framework that formulates knowledge distillation as a reinforcement learning (RL) policy optimization problem. Specifically, the student model’s multi-step denoising process is modeled as a sequential decision-making task, with sparse reward signals derived from alignment to teacher outputs; the Proximal Policy Optimization (PPO) algorithm is employed to optimize long-step denoising policies. Unlike conventional distillation paradigms—characterized by fixed step counts and layer-wise matching—our approach supports dynamic step scheduling and is model-agnostic, enabling general-purpose distillation. The reward design is inherently extensible. Experiments demonstrate that our method achieves superior generation quality over existing distillation approaches using only 4–8 inference steps, while exhibiting strong generalization and stability across multiple benchmark datasets.
This work addresses the performance degradation and lack of theoretical guarantees in existing approaches when distilling bidirectional video diffusion models into autoregressive ones, primarily due to architectural mismatches. The authors propose Causal Forcing, a novel method that leverages an autoregressive teacher model to guide the initialization of the ordinary differential equation (ODE) solver, effectively bridging the architectural gap between bidirectional and causal attention mechanisms. This approach provides the first theoretical resolution to the frame-level non-injectivity issue inherent in autoregressive distillation, ensuring invertibility of the flow map and circumventing performance loss caused by conditional expectation solutions. Built upon an ODE-based distillation framework and integrating autoregressive diffusion with causal attention, the method achieves state-of-the-art results, surpassing prior art by 19.3%, 8.7%, and 16.7% on Dynamic Degree, VisionReward, and Instruction Following metrics, respectively, enabling high-quality, real-time interactive video generation.
This work addresses the inefficiency of autoregressive video diffusion models, whose fixed denoising schedules often lead to either computational redundancy or insufficient refinement, hindering real-time generation. To overcome this limitation, the authors propose DSA, a confidence-guided dynamic computation framework that introduces adaptive step allocation into such models for the first time. DSA employs a lightweight confidence head to predict the reliability of denoising at each frame and jointly trains the generator and confidence head using a distribution-matching distillation objective. During inference, the model adaptively terminates or continues denoising based on predicted confidence, without requiring additional data or heuristic rules. Evaluated on an H100 GPU, the method achieves 22.63 FPS with sub-second latency while matching or surpassing state-of-the-art models in VBench quality metrics.
Existing distillation methods for video diffusion models directly adopt image-based distillation strategies, often leading to issues such as oversaturation, temporal inconsistency, and mode collapse. This work proposes the first distillation framework specifically designed for video diffusion models, featuring an adaptive regression loss that dynamically modulates spatial supervision strength and a temporal regularization loss to suppress inter-frame inconsistencies. Coupled with a frame interpolation strategy during inference, the method enables efficient, high-quality video generation at extremely low sampling steps. Evaluated on the VBench and VBench2 benchmarks, the proposed approach significantly outperforms existing distillation techniques, achieving superior perceptual fidelity, natural motion dynamics, and stable few-step synthesis.