🤖 AI Summary
This study addresses the limitation that weak convergence in constant-step-size stochastic approximation cannot guarantee accurate moment predictions. To overcome this, we construct a moment-exact Gaussian mixture model by matching stationary energy with local Ornstein–Uhlenbeck limits. Furthermore, Student-t noise modeling is introduced to capture secondary tail mass invisible to weak convergence, achieving uniform higher-order accuracy while accommodating singular covariances and non-limiting weight scenarios. The main contribution lies in quantifying second-order Wasserstein error bounds to obtain observable covariance and expected objective gaps. Numerical experiments confirm the framework’s value in precisely predicting SGD behavior across diverse geometries.
📝 Abstract
Local Gaussian models of constant-step learning predict output variability and expected losses, but weak convergence alone does not justify these moment predictions. We establish moment-accurate Gaussian mixtures by matching stationary energy with local Ornstein--Uhlenbeck limits, ruling out quadratic tail mass invisible to weak convergence. For step size $a$, the second-order Wasserstein error is $o(\sqrt a)$, uniformly over invariant laws, using each law's actual root weights. The assumptions combine confinement, descent, finitely many hyperbolic equilibria and root continuity with finite-variance innovations. The result yields observable covariances, expected objective gaps and first-order mean shifts, while allowing singular covariances, compatible saddles and weights without a limit. For additive noise given by a fixed invertible transform of independent standardized Student $t_3$ coordinates, symmetry gives an order-sharp $\sqrt a$ smooth-test bound. Numerical transport calculations demonstrate the value of root-specific covariances; controlled SGD studies assess observable predictions across step sizes, batch sizes and model geometries.