PDMD: Projected Distribution Matching Distillation for Video Diffusion Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses sample oversaturation and artifact issues arising from critic error accumulation in video diffusion model distillation by proposing Projected Distribution Matching Distillation (P-DMD). The method theoretically establishes that the residual constitutes an unbiased error estimate, and employs vector projection to filter erroneous components from gradient updates, thereby stabilizing training. P-DMD achieves efficient denoising in high-dimensional spaces while preserving informative signals, without requiring additional loss functions or architectural modifications. Notably, it can be integrated into mainstream video generation models such as Wan2.1 with a single line of code change. Experimental results demonstrate that P-DMD attains a VBench score of 83.73 under four-step evaluation, significantly outperforming baselines, with its superior audiovisual quality further corroborated by comprehensive user studies.
📝 Abstract
Modern video diffusion models require tens of denoising evaluations over long spatiotemporal token sequences. Distribution Matching Distillation (DMD) reduces the number of function evaluations (NFE) to just a few. However, DMD samples can degrade during training, exhibiting progressive oversaturation and artifacts. We trace this instability to critic errors, which enter successive student updates and accumulate over time. We introduce Projected Distribution Matching Distillation (PDMD) to filter critic errors. PDMD projects out the component of the DMD update parallel to the student-critic endpoint residual. At a fixed noisy query, we prove that this residual is an unbiased estimate of the critic's endpoint error. Under high-dimensional assumptions, this projection removes a constant fraction of critic error while discarding only a vanishing fraction of ideal DMD signal. Empirically, the projection stabilizes training and improves sample quality where DMD degrades and develops unnatural textures. PDMD requires only a one-line code change to DMD, with no extra loss, network, data, model pass, or multi-stage training. With Wan2.1, PDMD achieves a VBench total score of 83.73 at 4 NFE, surpassing matched DMD by 1.03 points. On MiniMax-H3 joint video-audio generation, PDMD achieves a VideoGen-Eval visual total score of 83.17, 0.41 points above the strongest distilled baseline. PDMD also achieves the best performance on all six audio metrics among the compared 4-NFE models. Qualitative comparisons and user studies favor PDMD over the distilled baselines in visual quality, motion, and audio quality. Code and models are available at https://pdmd2026.github.io/.
Problem

Research questions and friction points this paper is trying to address.

Video Diffusion Models
Distribution Matching Distillation
Critic Errors
Training Instability
Sample Degradation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Projected Distribution Matching Distillation
Video Diffusion Models
Critic Error Filtering
Few-step Generation
Distribution Matching Distillation
🔎 Similar Papers