🤖 AI Summary
This study addresses the limitation of existing distribution matching distillation (DMD) methods, which predominantly rely on latent diffusion models and overlook the characteristics and design potential of native RGB pixel space. To optimize the DMD interface for pixel-space teacher models, this work proposes a fixed high-noise matching band alongside an external visual representation guidance mechanism. Furthermore, it introduces DINO-Adv to eliminate discriminator gradient pathways and designs a parameter-free AF-Loss to facilitate semantic distribution alignment, thereby achieving efficient pixel-level supervision. Experimental results demonstrate that the proposed student model, requiring only four generation steps, outperforms both the 25-step teacher model and existing few-step distillation approaches across multiple benchmarks.
📝 Abstract
Distribution matching distillation (DMD) provides a general framework for few-step diffusion generation, but its modern text-to-image instantiations have been developed primarily around latent diffusion. It therefore overlooks key properties and design opportunities of native RGB. We revisit two DMD interfaces for pixel-space teachers. On the teacher-matching side, diagnostics show low-noise RGB matching is dominated by a local-texture cue, motivating a fixed high-noise matching band. On the real-data side, native clean-RGB outputs allow guidance from an external visual representation without traversing a decoder or sharing the heavy fake-score critic. DINO-Adv removes this critic from the adversarial gradient path and supplies local parametric patch guidance. For distribution-level guidance, we introduce AF-Loss, a parameter-free auxiliary semantic distribution-field objective designed for text-to-image DMD. It operates on detached rolling real and generated supports in the shared DINOv2 space while preserving prompt-conditioned teacher supervision. AF-Loss adds no learnable parameters or inference-time computation. Together these designs form DMA$^2$. Across DPG-Bench, GenEval, VQAScore, and COCO30K, the four-step DMA$^2$ student performs better than the 25-step teacher and evaluated few-step distillers.