🤖 AI Summary
This work investigates the long-term behavior and invariant distribution of constant-stepsize stochastic gradient descent (SGD) when optimizing objective functions possessing flat minima. Focusing on convex functions with local flatness exponent $m$, the analysis integrates Markovian noise modeling, Wasserstein contraction arguments, weak convergence to stochastic differential equations (SDEs), and asymptotic expansions of the local Hessian to establish that the SGD iterates converge geometrically to a unique invariant distribution. This limiting distribution is non-Gaussian, concentrates at scale $\alpha^{1/m}$, and converges weakly to the stationary distribution of the associated SDE. The study further identifies the contraction factor as $1 - c\alpha^{m-1}$ and extends the scaling law to anisotropic, separable settings.
📝 Abstract
For stochastic gradient descent (SGD) with a constant stepsize $α$, the invariant law of the iterates, centered at a minimizer, describes the behavior of the algorithm over long time horizons. In the strongly convex case, this invariant law has the familiar $\sqrtα$ scaling and a Gaussian limit as $α\downarrow 0$. We show that this behavior changes fundamentally for convex objectives $H$ with flat minima and (sub)quadratic tails.
More specifically, we study SGD with Markovian noise generated by a contractive driving chain. For every sufficiently small constant stepsize $α$, we prove existence, uniqueness, and geometric convergence to an augmented invariant law in a Wasserstein distance induced by an $α$-dependent metric. When the minimizer $x_\star$ has local flatness exponent $m\ge2$, meaning that $\nabla^2 H(x)\asymp \lVert x-x_\star\rVert^{m-2} I_d$ as $x\to x_\star$, we obtain a contraction bound with factor $1-cα^{m-1}$, where $c>0$ is a constant. This recovers the factor $1-cα$ in the quadratic case $m=2$. We then analyze the small-stepsize scaling limit. We show that the invariant law concentrates on the scale $α^{1/m}$ and that the rescaled iterates converge weakly to the stationary distribution of the stochastic differential equation $$
dY_t=-h_0(Y_t)\,dt+Σ^{1/2}\,dB_t , $$ where $h_0$ is the limiting drift at the minimizer and $Σ$ denotes the asymptotic covariance. This recovers the Gaussian limit when $m=2$ and gives generally non-Gaussian stationary limits in the flat case $m>2$. Finally, we give corresponding results for coordinate-separable objectives with unequal flatness exponents.