🤖 AI Summary
This study addresses the lack of rigorous theoretical guarantees for the asymptotic bias of Stochastic Gradient Langevin Dynamics (SGLD). Under strong convexity and Lipschitz conditions, it integrates stochastic differential equations, optimal transport, and convex optimization to derive exact upper bounds on the bias in Wasserstein distance, systematically analyzing how step size and moment assumptions affect convergence rates. The results demonstrate that the bias is $\mathcal{O}(h)$ under a fourth-moment assumption but degrades to $\mathcal{O}(h^{1/2})$ when only a second moment is available. Furthermore, a counterexample is constructed to confirm that the second moment alone is insufficient to ensure uniform first-order convergence. This work elucidates the fundamental limitations imposed by the moment conditions of noise distributions on SGLD convergence rates, filling a critical gap in the existing theory.
📝 Abstract
We prove asymptotic bias bounds for stochastic gradient Langevin dynamics in Wasserstein distance of order two. We assume that the negative log-density is strongly convex with a Lipschitz gradient, and that the stochastic gradient estimator is unbiased with an error satisfying a mean-square Lipschitz condition. The bounds are of order $h$ under a fourth moment assumption on the stochastic gradient error and of order $h^{1/2}$ under only a second moment assumption, where $h$ is the stepsize. A spiked-noise example shows that a second moment assumption alone is insufficient for a bound of order $h$ that is uniform over noise distributions with a fixed variance.