Adam under Generalized Smoothness with Second-Moment-Type Stochastic Gradients

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the long-standing challenge that convergence proofs for the Adam algorithm under generalized smooth objectives have heavily relied on strong tail assumptions, such as bounded or sub-Gaussian gradients. This work overcomes these limitations by demonstrating that second-moment information alone suffices to guarantee convergence. The theoretical analysis extends a self-normalization strategy to ensure iterates remain within well-behaved regions, and integrates stopping times, de-preconditioning, and L0-Lp generalized smoothness techniques. Consequently, this research establishes high-probability convergence rates for p<2, with a confidence dependence of δ^{-1/2} proven to be tight. These results fundamentally resolve the difficulty of analyzing Adam's convergence under weak moment conditions.
📝 Abstract
Adam is widely observed to remain stable even when the objective deviates significantly from global smoothness. Under the generalized smoothness framework, however, existing analyses rely on strong tail assumptions on the stochastic gradients, such as almost-sure boundedness or sub-Gaussianity. Whether Adam converges on generalized smooth objectives under only second moment information on the stochastic gradients, without such concentration assumptions, was identified as an important open direction by Li et al. (2023). This paper gives an affirmative answer under fairly general conditions: such tail assumptions are not necessary. Building on the Adam self-normalization framework of Jin et al. (2026), developed for classical smoothness and bounded variance, we extend the stopping-time and de-preconditioning strategy to the $L_0$-$L_p$ generalized smoothness condition and a generalized second moment ABC condition. Even when the stochastic-gradient condition provides only second moment information that may grow along the trajectory, the stochastic trajectory of Adam remains in a locally well-behaved smoothness region, with stretched-exponential tail decay under bounded variance and global smoothness. Consequently, we establish high-probability convergence rate guarantees over the full range $p<2$, with confidence dependence of order $δ^{-1/2}$, while the stepsize prefactor depends on $δ$ only through a single logarithmic factor. We further construct a hard instance showing that, under only second-moment information, this $δ^{-1/2}$-type confidence dependence is sharp. Finally, in the regime $p<1$, we combine the trajectory control with polynomial-growth estimates on rare events to obtain convergence rate guarantees in expectation.
Problem

Research questions and friction points this paper is trying to address.

Adam optimizer
generalized smoothness
stochastic gradients
second-moment condition
convergence analysis
💼 Related Jobs
No related jobs found.