🤖 AI Summary
This paper addresses the lack of a unified theoretical foundation for guidance mechanisms in diffusion models. We propose the first integrated algorithmic-theoretic framework unifying classifier-guided diffusion (CFG) and reward-guided diffusion. Methodologically, we introduce a reward-gradient term based on score differences into the reverse diffusion process, enabling efficient reward guidance without requiring full-trajectory reward modeling or training. Theoretically, we establish that CFG inherently optimizes the expected reciprocal of classifier probabilities and develop novel analytical tools—score-difference analysis and an upper bound estimator for expected reward—under our framework. Our contributions are fourfold: (1) a unified theoretical interpretation of both guidance paradigms; (2) a new sampling paradigm eliminating the need for trajectory-based reward training; (3) rigorous theoretical guarantees that our guidance strategy strictly improves reward metrics; and (4) empirical validation demonstrating superior generation quality and training stability across diverse benchmarks.
📝 Abstract
Guided or controlled data generation with diffusion modelslfootnote{Partial preliminary results of this work appeared in International Conference on Machine Learning 2025 citep{li2025provable}.} has become a cornerstone of modern generative modeling. Despite substantial advances in diffusion model theory, the theoretical understanding of guided diffusion samplers remains severely limited. We make progress by developing a unified algorithmic and theoretical framework that accommodates both diffusion guidance and reward-guided diffusion. Aimed at fine-tuning diffusion models to improve certain rewards, we propose injecting a reward guidance term -- constructed from the difference between the original and reward-reweighted scores -- into the backward diffusion process, and rigorously quantify the resulting reward improvement over the unguided counterpart. As a key application, our framework shows that classifier-free guidance (CFG) decreases the expected reciprocal of the classifier probability, providing the first theoretical characterization of the specific performance metric that CFG improves for general target distributions. When applied to reward-guided diffusion, our framework yields a new sampler that is easy-to-train and requires no full diffusion trajectories during training. Numerical experiments further corroborate our theoretical findings.