🤖 AI Summary
This work addresses the limitation of classical stochastic optimization theory, which relies on the uniform escape (UE) assumption to avoid strict saddle points—a condition often violated in over-parameterized, interpolation, or finite-sum settings. The authors establish a stochastic recursive almost sure saddle avoidance theorem without requiring the UE assumption. By introducing a path-dependent variable transformation and a pathwise Lyapunov–Perron method, they extend the center-stable manifold framework to sequences of random mappings lacking common fixed points, under assumptions of local smoothness, finite moment conditions, and without-replacement sampling structure. This unified framework applies to stochastic mirror descent (including SGD), stochastic reshuffling, and proximal stochastic gradient methods for nonsmooth composite objectives, proving their almost sure convergence to local minima by first avoiding strict saddle points and then ensuring iterative convergence.
📝 Abstract
Unit excitation (UE) is a common assumption in stochastic saddle avoidance: the stochastic error must have a uniformly positive component along every direction, in expectation. This condition gives a direct way to rule out convergence to strict saddles, but it also oversimplifies the actual noise structure, and does not match many stochastic optimization regimes. In overparameterized or interpolation models, the noise may vanish near stationarity. In finite-sum problems, the stochastic gradient noise may lie in a low-dimensional, data-dependent subspace. In these (common) scenarios, UE is naturally not satisfied. In this paper, we prove an abstract almost sure avoidance theorem for stochastic recursions without UE. The theorem replaces UE-type requirements by verifiable pathwise conditions. In applications, these conditions follow, e.g., from local smoothness and finite-moment assumptions under standard i.i.d. sampling, or from the finite-sum structure under without-replacement sampling. Since the stochastically sampled maps generally do not share a fixed point, the celebrated center-stable manifold argument used in deterministic analyses is not directly applicable. Instead, we use a path-dependent change of variables together with a pathwise Lyapunov--Perron-based proof strategy. As applications, we obtain strict saddle avoidance for stochastic mirror descent (including SGD) and for random reshuffling. For nonsmooth composite objectives, we prove avoidance results for a proximal-type stochastic gradient method. Combining these insights with suitable iterate convergence guarantees, this allows establishing convergence to local minimizers of the original objective function.