🤖 AI Summary
This study addresses the high inference costs of diffusion models and the over-smoothing artifacts commonly introduced by single-step distillation. To overcome these limitations, this work proposes a staged velocity distillation method that decouples the generation process into coarse and fine stages. Departing from the conventional monolithic student model paradigm, each stage is independently handled by a half-sized expert network dedicated to structure and detail optimization, respectively. By integrating staged average velocity modeling with diffusion distillation, the total computational cost is reduced to that of a single full-network forward pass. Evaluated on ImageNet, the proposed approach achieves an FID of 1.48, significantly outperforming existing distillation methods while reducing active parameters and peak memory consumption by approximately 50%, thereby enabling high-quality, low-cost, and rapid image generation.
📝 Abstract
Recent diffusion-based image generation backbones have grown substantially in scale, making the network inference cost increase rapidly. While diffusion distillation techniques can reduce the number of inference steps, high-quality image generation within a single full-backbone-forward compute budget remains challenging. Existing one-step methods typically allocate this budget to a single evaluation of a monolithic student. However, approximating the heterogeneous coarse-to-fine transport with a single monolithic mapping is difficult and often leads to over-smoothed outputs. To address this issue, we propose Phase-wise Velocity Distillation (PVD), which partitions the generation timeline into a coarse and a fine phase, and models the transition within each phase via the average velocity. A dedicated half-sized expert is assigned to each phase, decoupling structural composition from detail refinement while keeping the cumulative computation equivalent to one full-backbone forward pass. We show that the use of two half-sized phase-specific experts outperforms a single full-size monolithic student. On class-conditional image generation, PVD achieves an FID of 1.48 on ImageNet 256 x 256. On more complex text-to-image (T2I) tasks, PVD-distilled models (Stable Diffusion 3.5-Medium, FLUX.1-dev, Qwen-Image) produce results competitive with their multi-step teachers, significantly outperforming prior distillation methods. Moreover, across the evaluated T2I backbones, PVD reduces active parameters by 49.10-50.89% and peak VRAM by 45.76-48.36% compared to the corresponding teachers. Source code and distilled models are available at https://github.com/PolyU-VCLab/PVD.