DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

๐Ÿ“… 2026-10-01
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the additional computational and memory overhead incurred by maintaining auxiliary diffusion models in distribution matching distillation. It reformulates distribution matching as a classification task and proposes a log-density ratio estimation method based on a shared backbone discriminator head. By introducing a gap reweighting mechanism to adaptively modulate teacher supervision, we theoretically prove that the optimal discriminator solution precisely recovers the distribution matching gradient. This approach integrates adversarial distillation with linear loss optimization to achieve efficient few-step visual generation. Extensive experiments demonstrate state-of-the-art performance across ImageNet, COCO, and video benchmarks, with human preference evaluations significantly outperforming DMD2 and rCM.
๐Ÿ“ Abstract
Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the student without auxiliary score fitting. We prove that at the discriminator optimum these losses recover the distribution-matching gradient underlying DMD, through the classical identity linking discriminator logits to log-density ratios. We further introduce gap-based reweighting, which adapts teacher supervision across noise levels from the real-data head's empirical logit gap between real and teacher samples. DMAD reaches a Frรฉchet Inception Distance (FID) of 1.04 with one-step generation on ImageNet-64x64, 14.47 with four-step SDXL on COCO-10K, and a VBench total score of 85.15 with four-step Wan2.1-T2V-14B, the best values among the compared few-step methods and the multi-step teachers. On MiniMax-H3-33B, our four-step student achieves overall human preference rates of 79.1% over DMD2 and 84.6% over rCM for joint audio-video generation, excluding ties. Our code, models and demos are available at https://yzmblog.github.io/projects/DMAD.
Problem

Research questions and friction points this paper is trying to address.

Distribution Matching Distillation
Adversarial Distillation
Few-step Generation
Visual Generation
Auxiliary Model Overhead
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial Distillation
Distribution Matching
Discriminator Logits
Gap-based Reweighting
Few-step Generation
๐Ÿ”Ž Similar Papers
No similar papers found.