GPD: Guided Progressive Distillation for Fast and High-Quality Video Generation

📅 2026-02-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the low sampling efficiency of video diffusion models, where existing acceleration methods often compromise generation quality. The authors propose a guided progressive distillation framework in which a teacher model progressively instructs a student model to denoise with larger step sizes. To balance computational efficiency and detail preservation while reducing optimization difficulty, the approach integrates online target generation and latent-space frequency-domain constraints. By unifying diffusion distillation, online target synthesis, and frequency-domain regularization, the method reduces the sampling steps of the Wan2.1 model from 48 to 6 while maintaining competitive visual quality on the VBench benchmark, significantly outperforming current distillation techniques.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Deep Generative Models & AutoencodersSearch and Optimization: Sampling/Simulation-based Search

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Diffusion models have achieved remarkable success in video generation; however, the high computational cost of the denoising process remains a major bottleneck. Existing approaches have shown promise in reducing the number of diffusion steps, but they often suffer from significant quality degradation when applied to video generation. We propose Guided Progressive Distillation (GPD), a framework that accelerates the diffusion process for fast and high-quality video generation. GPD introduces a novel training strategy in which a teacher model progressively guides a student model to operate with larger step sizes. The framework consists of two key components: (1) an online-generated training target that reduces optimization difficulty while improving computational efficiency, and (2) frequency-domain constraints in the latent space that promote the preservation of fine-grained details and temporal dynamics. Applied to the Wan2.1 model, GPD reduces the number of sampling steps from 48 to 6 while maintaining competitive visual quality on VBench. Compared with existing distillation methods, GPD demonstrates clear advantages in both pipeline simplicity and quality preservation.
Problem

Research questions and friction points this paper is trying to address.

video generation
diffusion models
computational cost
quality degradation
sampling steps
Innovation

Methods, ideas, or system contributions that make the work stand out.

Guided Progressive Distillation
video generation
diffusion acceleration
frequency-domain constraints
knowledge distillation