🤖 AI Summary
Existing work lacks a theoretical explanation for why diffusion guidance consistently yields high-quality samples, particularly regarding how it ensures generated samples remain within the support of the target distribution. This work introduces the notion of “support robustness” and, under the assumption of an exact score function, rigorously proves for the first time that diffusion guidance almost surely produces samples arbitrarily close to the target support, thereby guaranteeing structural plausibility and physical meaningfulness. The analysis unifies a broad class of discretization schemes—including DDPM, DDIM, and those based on exponential integrators—providing a solid theoretical foundation for the empirical success of diffusion guidance in generating high-fidelity samples.
📝 Abstract
Diffusion guidance is a powerful technique that enables controllable and high-fidelity sample generation with diffusion models. At a high level, it modifies the score function by incorporating a guidance term that steers the generative process toward a desired condition. Despite its empirical success, the theoretical properties of diffusion guidance remain largely unexplored, and it is not well understood why it consistently produces high-quality samples.
In this work, we explain the effectiveness of diffusion guidance by establishing a \emph{robustness of support} property. Specifically, we show that, given exact access to the score functions, guided diffusion processes almost always generate samples that remain close to the target support. This property is particularly desirable, as samples that lie off the support are often structurally implausible and may adversely affect downstream tasks. Our analysis covers both Denoising Diffusion Implicit Models (DDIM) and Denoising Diffusion Probabilistic Models (DDPM), and applies to a wide range of discretization schemes induced by exponential integrators. Our results provide a rigorous foundation for understanding why diffusion guidance produces physically meaningful and structurally plausible samples.