🤖 AI Summary
This work investigates the long-term approximation accuracy of stochastic gradient Langevin dynamics (SGLD) to continuous Langevin diffusion, focusing on uniform-in-time error bounds for the Kullback–Leibler (KL) divergence and Wasserstein/total variation distances between their invariant measures. Leveraging a synthesis of stochastic differential equation analysis, information-theoretic entropy estimation, and diffusion approximation theory under non-convex potentials, we establish, for the first time, a sharp, uniform-in-time $O(eta^2)$ upper bound on the KL divergence for step size $eta$. This directly implies $O(eta)$ bounds on the Wasserstein and total variation distances between the invariant measures. The results hold for general non-convex potentials and accommodate variable step sizes—significantly improving upon prior $O(eta)$ KL bounds. To date, this provides the strongest theoretical guarantee for the stability and statistical fidelity of SGLD in Bayesian inference and sampling.
📝 Abstract
We establish a sharp uniform-in-time error estimate for the Stochastic Gradient Langevin Dynamics (SGLD), which is a widely-used sampling algorithm. Under mild assumptions, we obtain a uniform-in-time $O(eta^2)$ bound for the KL-divergence between the SGLD iteration and the Langevin diffusion, where $eta$ is the step size (or learning rate). Our analysis is also valid for varying step sizes. Consequently, we are able to derive an $O(eta)$ bound for the distance between the invariant measures of the SGLD iteration and the Langevin diffusion, in terms of Wasserstein or total variation distances. Our result can be viewed as a significant improvement compared with existing analysis for SGLD in related literature.