GFD-OPD: Guidance-Folded On-Policy Distillation of Diffusion Models Across Scales

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of standard online distillation methods from large to small diffusion models, where classifier-free guidance (CFG) amplifies distributional discrepancies between teacher and student. To overcome this challenge, we propose the GFD-OPD framework. We first introduce a novel Fixed-State KL divergence metric that reveals the error accumulation mechanism induced by CFG across cross-scale models. Subsequently, we design a guidance folding strategy that effectively narrows the teacher-student gap and suppresses error propagation, thereby enabling efficient online policy distillation. Extensive experiments demonstrate that our approach significantly improves both training efficiency and generation quality, achieving state-of-the-art performance across all evaluated benchmarks.
📝 Abstract
On-policy distillation (OPD) has demonstrated two important capabilities in language models: compressing large teachers into smaller students and merging expert models into a single model. Existing diffusion OPD, however, mostly focus on the latter, with teachers and students sharing the same backbone and scale. We investigate large-to-small diffusion opd from large teachers to a small student and find that the standard recipe fails. To find the underlying cause, we propose Fixed-State KL, an effective and fair way to measure the distribution gap between student and teacher during OPD training for diffusion models. We are the first to clarify why large-to-small OPD is challenging for diffusion models: a smaller student struggles to perfectly match the distribution of a larger teacher, while classifier-free guidance can accumulate and amplify the distributional discrepancies between the student's conditional and unconditional branches and those of the teacher. To solve this problem, we propose GFD-OPD, a simple yet effective method that reduces the student-teacher gap while avoiding the error amplification of the CFG composition. Across numerous experiments, GFD outperforms previous baselines in both training efficiency and final performance, achieving state-of-the-art results on all benchmarks.
Problem

Research questions and friction points this paper is trying to address.

On-policy distillation
Diffusion models
Large-to-small distillation
Classifier-free guidance
Distribution gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Diffusion Models
Classifier-Free Guidance
Fixed-State KL
Large-to-Small Distillation
🔎 Similar Papers
No similar papers found.