Optimal Projection-Free Adaptive SGD for Matrix Optimization

📅 2026-04-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing projection-free adaptive SGD methods in matrix optimization, which typically require additional hyperparameter tuning and lack dimension-independent convergence guarantees. By analyzing the stability of the Leon preconditioner, we propose a practical one-sided Shampoo algorithm that eliminates the need for per-iteration projections and extra hyperparameter tuning, and for the first time incorporates Nesterov acceleration. Leveraging block-diagonal preconditioning and online convex optimization techniques, we develop a unified analytical framework for nonsmooth nonconvex optimization, achieving dimension-independent convergence rates. This approach significantly enhances optimization performance while maintaining computational efficiency.

Technology Category

Search and Optimization: Non-convex OptimizationMachine Learning: OptimizationNatural Language Processing: Learning & Optimization for NLP

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationResponsible Web: Human-perceived consequences of algorithmic deployment on the web
📝 Abstract
Recently, Jiang et al. [2026] developed Leon, a practical variant of One-sided Shampoo [Xie et al., 2025a, An et al., 2025] algorithm for online convex optimization, which does not require computing a costly quadratic projection at each iteration. Unfortunately, according to the existing analysis, Leon requires tuning an additional hyperparameter in its preconditioner and cannot achieve dimension-independent convergence guarantees for convex optimization problems beyond the bounded gradients assumption. In this paper, we resolve this issue by proving certain stability properties of Leon's preconditioner. Using our improved analysis, we show that tuning the extra hyperparameter can be avoided and, more importantly, develop the first practical variant of One-sided Shampoo with Nesterov acceleration, which does not require computing projections at each iteration. As a side contribution, we obtain improved dimension-independent rates in the non-smooth non-convex setting and develop a unified analysis of the proposed algorithm, which yields accelerated projection-free adaptive SGD with (block-)diagonal preconditioners.
Problem

Research questions and friction points this paper is trying to address.

matrix optimization
projection-free
adaptive SGD
dimension-independent convergence
hyperparameter tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

projection-free
adaptive SGD
Nesterov acceleration
preconditioner stability
dimension-independent convergence