🤖 AI Summary
This study addresses the theoretical gap in analyzing the convergence rate of the last iterate for normalized gradient descent in Hölder smooth convex optimization. By leveraging the Performance Estimation Problem (PEP) framework, we reveal the inherent logarithmic overhead associated with constant step sizes and establish corresponding convergence upper bounds. Furthermore, we propose a linearly decaying step size strategy that requires no prior knowledge, supported by systematic numerical validation and theoretical analysis. Our results demonstrate that this strategy matches optimal iteration complexity, achieving a last-iterate convergence rate of $O(T^{-(1+\nu)/2})$. By successfully eliminating the extraneous logarithmic factor, this work significantly enhances the practical convergence efficiency of the algorithm during deployment.
📝 Abstract
Normalized gradient descent is a widely studied adaptive optimization method. Most existing analyses focus on the best iterate or a weighted average of the iterates, whereas practical implementations typically return the last iterate. In this paper, we study the last-iterate convergence of normalized gradient descent for convex, $(\nu,M_\nu)$-H\"older-smooth objectives. For a constant stepsize, we establish an upper bound of $\mathcal{O}\bigl((\log^2(T)/T)^{(1+\nu)/2}\bigr)$, which contains a logarithmic overhead relative to the known $\mathcal{O}\bigl(T^{-(1+\nu)/2}\bigr)$ guarantees for the best and weighted-average iterates. For $\nu = 0$, this overhead is known to be unavoidable. We complement this analysis with numerical results based on the performance estimation problem (PEP), investigating the finite-horizon worst-case behavior in the smooth setting and whether the logarithmic overhead reflects an intrinsic limitation of constant-step normalized gradient descent. We then show that a linearly decreasing stepsize yields a last-iterate guarantee of $\mathcal{O}\bigl(T^{-(1+\nu)/2}\bigr)$, matching the order of the best-iterate/weighted-average guarantees without requiring knowledge of $\nu$ and $M_\nu$.