Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size

📅 2025-01-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the challenges of gradient computation bottlenecks and limited computational resources in large-scale Riemannian optimization. To accelerate stochastic optimization convergence, we propose and theoretically analyze an incremental batch-size strategy. We establish, for the first time on Riemannian manifolds, the optimal $O(1/sqrt{T})$ convergence rate for Riemannian Stochastic Gradient Descent (RSGD), improving upon the prior best-known bound of $O(sqrt{log T}/T^{1/4})$. Our method integrates dynamic batch-size scheduling—employing either polynomial or exponential growth—cosine annealing, and polynomial learning-rate decay. Empirical evaluation on Principal Component Analysis (PCA) and low-rank matrix completion tasks demonstrates consistent superiority of the incremental batch strategy over fixed batch sizes in both convergence speed and final accuracy—except on MovieLens—particularly benefiting large-scale models and resource-constrained environments.

Technology Category

Machine Learning: Learning with ManifoldsReasoning under Uncertainty: Stochastic OptimizationSearch and Optimization: Non-convex Optimization

Application Category

Search and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSecurity and Privacy: Large-scale security measurements
📝 Abstract
Many models used in machine learning have become so large that even computer computation of the full gradient of the loss function is impractical. This has made it necessary to efficiently train models using limited available information, such as batch size and learning rate. We have theoretically analyzed the use of Riemannian stochastic gradient descent (RSGD) and found that using an increasing batch size leads to faster RSGD convergence than using a constant batch size not only with a constant learning rate but also with a decaying learning rate, such as cosine annealing decay and polynomial decay. In particular, RSGD has a better convergence rate $O(frac{1}{sqrt{T}})$ than the existing rate $O(frac{sqrt{log T}}{sqrt[4]{T}})$ with a diminishing learning rate, where $T$ is the number of iterations. The results of experiments on principal component analysis and low-rank matrix completion problems confirmed that, except for the MovieLens dataset and a constant learning rate, using a polynomial growth batch size or an exponential growth batch size results in better performance than using a constant batch size.
Problem

Research questions and friction points this paper is trying to address.

Machine Learning
Model Training Acceleration
Large-scale Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Riemann Stochastic Gradient Descent
Batch Size Increment
Learning Acceleration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kanata Oowada
Department of Computer Science, Meiji University, Japan
H
Hideaki Iiduka
Department of Computer Science, Meiji University, Japan