🤖 AI Summary
This study investigates the convergence behavior of AdaGrad in stochastic convex optimization, addressing the open theoretical challenges concerning its last-iterate rate and high-probability averaged-iterate bounds. Through sub-Gaussian noise analysis and tight lower bound construction, we demonstrate that no universal last-iterate convergence rate exists for AdaGrad and that bounded variance alone is insufficient to guarantee high-probability convergence. Furthermore, this work establishes for the first time the inevitability of the logarithmic factor present in existing upper bounds, while eliminating this penalty under general ABC conditions to achieve the optimal O(1/√T) convergence rate. By revealing the theoretical limitations and necessary conditions for the convergence of AdaGrad, this research provides a solid theoretical foundation for the design of adaptive optimization algorithms.
📝 Abstract
AdaGrad and AdaGrad-Norm are widely used adaptive methods, but their precise behavior in stochastic convex optimization remains less understood. We first show that AdaGrad-Norm and AdaGrad do not admit any universal \textbf{last-iterate} rate, even under sub-Gaussian noise and bounded iterates. We then prove that bounded variance alone is too weak: even with bounded iterates, it cannot yield \textbf{high-probability average-iterate} rates, and without bounded iterates it may not even guarantee convergence in expectation. Besides, we construct tight $\Omega(\log T/\sqrt T)$ lower bounds for both AdaGrad-Norm and AdaGrad under sub-Gaussian noise, showing that $\log T$ in existing average-iterate upper bounds is unavoidable. Finally, we show that this $\log T$ loss disappears once bounded-iterate condition is imposed: under a general ABC condition, both methods achieve rates ${O}(1/\sqrt{T})$.