rgd convergence analysis

Analyzes Riemannian gradient descent algorithms to establish their convergence behavior by proving iteration-complexity and rate bounds (e.g., O(log(1/ε)), linear, or sublinear rates) and conditions for reaching ε-optimality. Constructs mathematical proofs that bound final suboptimality and characterize its dependence on manifold geometry, step sizes, initialization, stochasticity, and model misspecification, and adapts these analyses to algorithmic variants (stochastic, constrained, or approximate updates).

rgdconvergenceanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Finite-Time Analysis of Stochastic Nonconvex Nonsmooth Optimization on the Riemannian Manifolds

Oct 24, 2025
ES
Emre Sahinoglu
🏛️ Northeastern University | Tsinghua University

This work addresses the finite-time convergence of nonsmooth, nonconvex stochastic optimization on Riemannian manifolds—a setting lacking prior theoretical guarantees. To bridge this gap, we introduce, for the first time, a manifold-adapted Goldstein stationarity measure and propose two algorithms: the first-order RO²NC and the zero-order ZO-RO²NC. Both algorithms are proven to converge to a (δ, ε)-Goldstein stationary point with the optimal sample complexity of O(ε⁻³δ⁻¹), matching the best-known rate in Euclidean space. This constitutes the first finite-time convergence guarantee for fully nonsmooth, nonconvex stochastic optimization on Riemannian manifolds. Empirical evaluation demonstrates the efficacy and robustness of our methods on tasks including principal component analysis and manifold-constrained sparse regression.

Analyzes stochastic nonsmooth nonconvex optimization on Riemannian manifoldsDevelops zeroth-order method when gradient information is unavailableProposes RO2NC algorithm with finite-time convergence guarantees

Last-Iterate Convergence of Adaptive Riemannian Gradient Descent for Equilibrium Computation

Jun 29, 2023
YC
Yang Cai
🏛️ Yale University | UC Berkeley | Columbia University | University of Wisconsin-Madison

This work addresses the linear convergence of the last iterate in computing Nash equilibria of games defined on Riemannian manifolds, focusing on geodesically strongly monotone games and Riemannian gradient descent (RGD). Methodologically, it integrates tools from Riemannian optimization, geodesic monotonicity analysis, and game theory. The contributions are threefold: (i) it establishes the first geometry-agnostic linear convergence guarantee for the last iterate of RGD—i.e., independent of prior knowledge of manifold curvature; (ii) it proposes FARGD, an adaptive algorithm that achieves a convergence rate matching that of optimal non-adaptive methods without requiring estimates of the condition number; and (iii) it designs stochastic RGD (SRGD), attaining optimal sample complexity under stochastic gradient noise. Collectively, these advances significantly enhance the robustness and practical applicability of equilibrium computation on Riemannian manifolds.

Analyzing Riemannian gradient descent convergence on curved manifoldsEstablishing geometry-agnostic linear convergence for geodesic monotone gamesExtending convergence guarantees to stochastic and adaptive variants

This paper addresses the challenges of gradient computation bottlenecks and limited computational resources in large-scale Riemannian optimization. To accelerate stochastic optimization convergence, we propose and theoretically analyze an incremental batch-size strategy. We establish, for the first time on Riemannian manifolds, the optimal $O(1/sqrt{T})$ convergence rate for Riemannian Stochastic Gradient Descent (RSGD), improving upon the prior best-known bound of $O(sqrt{log T}/T^{1/4})$. Our method integrates dynamic batch-size scheduling—employing either polynomial or exponential growth—cosine annealing, and polynomial learning-rate decay. Empirical evaluation on Principal Component Analysis (PCA) and low-rank matrix completion tasks demonstrates consistent superiority of the incremental batch strategy over fixed batch sizes in both convergence speed and final accuracy—except on MovieLens—particularly benefiting large-scale models and resource-constrained environments.

Large-scale ModelsMachine LearningModel Training Acceleration

Geometry, Computation, and Optimality in Stochastic Optimization

Sep 23, 2019
CC
Chen Cheng
🏛️ Stanford University

This work systematically uncovers the decisive role of problem geometry—specifically, the curvature of the constraint set and the structure of gradients—in governing the statistical-computational trade-offs of stochastic and online optimization algorithms. We introduce the first geometric measure quantifying the deviation of a constraint set from quadratic convexity, rigorously identifying the geometric origins of suboptimality in subgradient methods. We prove that diagonal-preconditioned SGD achieves minimax-optimal convergence rates under quadratic convex constraints. For non-Euclidean, non-quadratically-convex domains—such as ℓₚ-balls with p < 2—we establish tight convergence bounds for mirror descent and adaptive gradient methods, and uncover, for the first time, a precise correspondence between their convergence rates and the accuracy-computation trade-off in Gaussian sequence estimation. Our results provide geometric criteria for algorithm selection and unify the understanding of when nonlinear updates—e.g., via mirror descent—are necessary to attain statistical optimality.

Characterize optimality of stochastic gradient methods via geometryDetermine when nonlinear updates are necessary for optimal convergenceQuantify sub-optimality of subgradient methods using constraint convexity

Directional Smoothness and Gradient Methods: Convergence and Adaptivity

Mar 06, 2024
AM
Aaron Mishkin
🏛️ Stanford University | Princeton University | Meta AI | Flatiron Institute

This work addresses the slow convergence of gradient descent on complex objectives and its reliance on strong global smoothness assumptions. We introduce *directional smoothness*, a novel geometric concept characterizing the local smoothness of the objective function along the optimization trajectory—thereby circumventing restrictive global Lipschitz continuity requirements. Leveraging this path-dependent characterization, we derive a trajectory-aware suboptimality bound and formulate an implicit adaptive step-size equation. We theoretically establish that Polyak’s step size and normalized gradient descent inherently achieve path-adaptive fast convergence. Our methodology integrates directional smoothness analysis, implicit step-size design, and convergence theory for both convex and nonconvex settings. Experiments on logistic regression demonstrate that our new bound substantially improves upon classical $L$-smoothness-based guarantees. Notably, this is the first work to provide path-dependent convergence rates for these two canonical algorithms without requiring prior knowledge of smoothness parameters.

Complex FunctionGradient DescentOptimization Efficiency

Latest Papers

What's happening recently
View more

This work addresses the gap between abstract Riemannian geometry and practical algorithmic implementation by systematically developing a computationally tractable geometric framework for Riemannian optimization. Focusing on canonical matrix manifolds—Stiefel, Grassmann, and symmetric positive-definite (SPD) manifolds—it explicitly derives core geometric structures, including tangent spaces, metric tensors, Levi-Civita connections, curvature operators, and geodesics, all expressed in coordinate- and matrix-based forms amenable to numerical computation. Furthermore, it provides closed-form expressions for the Riemannian gradient, Hessian, exponential map, and retraction operators. To the best of our knowledge, this is the first unified formulation that translates classical differential-geometric constructions into a consistent, implementation-ready framework, thereby bridging theory and practice and offering a rigorous foundation for efficient and accurate algorithm design in Riemannian optimization and geometric machine learning.

coordinate-level derivationsdifferential geometryimplementation gap

This work investigates whether gradient descent algorithms relying solely on predetermined non-negative stepsize schedules can achieve the optimal $O(T^{-2})$ last-iterate convergence rate in smooth convex optimization. By constructing adversarial instances and employing a refined recursive analysis, the authors establish—for the first time—a lower bound of $\Omega(T^{-1.9319})$ for this class of methods, rigorously demonstrating that stepsize scheduling alone is insufficient to attain the $O(T^{-2})$ rate. This result delineates the fundamental theoretical limitations of stepsize-scheduled gradient descent and fills a critical gap in the lower-bound analysis for such algorithms.

accelerationconvergence rategradient descent

This work proposes a trajectory-restricted framework for linear convergence analysis that overcomes the conservatism of traditional first-order methods, whose guarantees often rely on global geometric conditions and worst-case constants. Instead of imposing regularity assumptions globally, our approach requires only local geometric properties—such as restricted Polyak–Łojasiewicz inequalities, error bounds, and quadratic growth—on the subset of the space actually traversed by the algorithm. We establish explicit relationships among the associated constants and show that, for piecewise polyhedral composite problems, once iterates enter a well-conditioned active manifold, convergence is governed by the restricted Hoffman constant of that manifold, yielding an improved effective condition number and faster local convergence. The results demonstrate that linear convergence fundamentally depends on the local geometry encountered along the algorithmic trajectory, rather than on global worst-case scenarios.

geometric regularityHoffman constantlinear convergence

This work addresses the challenges of poor scalability, limited parallelizability, and complex subproblems in nonsmooth nonconvex optimization with orthogonality constraints by proposing a retraction-free primal-dual linearized smoothed augmented Lagrangian method. The proposed algorithm introduces, for the first time, a retraction-free primal-dual framework to orthogonality-constrained optimization, eliminating nested loops and intricate subproblem solvers in favor of a single-loop iteration scheme. Leveraging the Kurdyka–Łojasiewicz property, the method is theoretically shown to converge to an $\varepsilon$-KKT point with an iteration complexity of $O(\varepsilon^{-3})$, without requiring Riemannian retractions. Numerical experiments demonstrate that the algorithm significantly outperforms existing approaches in both computational efficiency and scalability.

nonconvex optimizationnonsmooth optimizationorthogonality constraints

Hot Scholars

YY

Yingzhen Yang

Arizona State University
Statistical Machine Learning and Deep Learning