Ordinary Nonconvex SGD under Distance-Dependent Moments: Finite-Horizon Stationarity and Nagaev Bounds

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that stochastic gradient variance in non-convex optimization varies with iteration distance, rendering traditional uniform moment assumptions inadequate. To overcome this, the authors analyze vanilla SGD under distance-dependent moment conditions by introducing a direct descent displacement argument framework. This approach integrates the Fuk-Nagaev inequality in Hilbert spaces with predictable localization techniques, eliminating the need for additional modifications such as gradient clipping, momentum, or batch size enlargement. The work demonstrates that relying solely on second-order moments suffices to achieve optimal complexity matching the Blum-Gladyshev lower bound. Furthermore, it establishes an average squared gradient stationary convergence rate of T^{-1/3} along with high-probability Nagaev-type bounds. These results theoretically confirm the optimality of unmodified SGD for this class of problems.
📝 Abstract
Uniform noise-moment bounds exclude stochastic gradients whose variability increases with the iterate. We study ordinary, single-sample stochastic gradient descent for smooth, lower-bounded, possibly nonconvex objectives under distance-dependent conditional moments. Under second moments alone, a direct descent--displacement argument yields $T^{-1/3}$ expected average squared-gradient stationarity with a horizon-dependent stepsize. An explicit oracle-complexity corollary matches the known smooth Blum--Gladyshev (BG-0) lower bound, including the $Lb_2Δ^3\varepsilon^{-6}$ and $LΔσ^2\varepsilon^{-4}$ stochastic terms, where $Δ$ is the initial objective gap and $σ^2+b_2\|x-x_1\|^2$ bounds the variance. Thus unchanged SGD attains the minimax stochastic complexity in this second-moment class. For $p>2$, predictable localization and a Hilbert-space Fuk--Nagaev inequality yield a high-probability bound separating logarithmic variance and polynomial rare-shock contributions. The localization radius is derived from the recursion: no bounded-iterate assumption, clipping, normalization, momentum, or increasing batch size is needed. We also give increasing-confidence rates, an objective-gap-growth refinement recovering root-$T$ stationarity, and stochastic $L^p$-Lipschitz examples. The broad BG-0 optimality statement is distinguished from the smaller mean-square-smooth class, in which additional oracle structure permits faster algorithms.
Problem

Research questions and friction points this paper is trying to address.

stochastic gradient descent
nonconvex optimization
distance-dependent moments
oracle complexity
high-probability bounds
Innovation

Methods, ideas, or system contributions that make the work stand out.

Stochastic Gradient Descent
Distance-Dependent Moments
Fuk-Nagaev Inequality
Oracle Complexity
Nonconvex Optimization
💼 Related Jobs
No related jobs found.