🤖 AI Summary
This study addresses the limitation that classical Langevin samplers are constrained by the condition number, a bottleneck whose circumvention by diffusion models has lacked rigorous theoretical justification. By leveraging Gaussian measure concentration and spectral bound analysis combined with first-order asymptotic matching and gradient descent estimation techniques, this work provides the first rigorous proof that time-dependent score trajectories eliminate condition number dependence during the sampling phase, clarifying that the advantage of noise injection does not originate from the learning stage. Furthermore, it establishes tight 2-Wasserstein convergence bounds under optimized hyperparameters, revealing the quantitative relationship between sampling error, dimensionality, step count, and covariance eigenvalues. The derived sampling error for diffusion processes is O(√(dλ_max)logN/N), significantly outperforming Langevin dynamics which suffer from an additional √κ factor.
📝 Abstract
Despite their empirical success, why diffusion models overcome the bottlenecks of classical score-based samplers remains unclear. In this work, we leverage Gaussian distributions to isolate this phenomenon. We establish 2-Wasserstein convergence bounds for optimized hyperparameters, showing that diffusion processes achieve a sampling error of $O(\sqrt{dλ_{\max}}\log N/N)$, where $d$ is the dimension, $N$ the number of sampling steps, and $λ_{\max}$ the largest eigenvalue of the target covariance matrix. Unadjusted and underdamped Langevin dynamics suffer from an additional $\sqrtκ$ factor, where $κ$ is the condition number. These rates follow from spectral bounds which are sharp: we confirm them via matching first-order asymptotics as $N\rightarrow\infty$. Our analysis provides a rigorous characterization, in the Gaussian setting, of how time-dependent score trajectories remove condition-number dependence during sampling. By contrast, in the learning phase, we show that estimating the unnoised score by gradient descent leads to essentially the same estimator as estimating a noisy score, which suggests that the benefits of noising do not come from the learning phase.