Unified convergence analysis for adaptive optimization with moving average estimator

📅 2021-04-30
🏛️ Machine-mediated learning
📈 Citations: 14
✨ Influential: 3
📄 PDF
🤖 AI Summary
This work addresses the weak theoretical convergence guarantees of adaptive optimization algorithms—particularly those employing moving-average momentum estimators—in nonconvex optimization. Methodologically, it establishes a unified convergence analysis framework: (i) it provides the first rigorous proof that monotonic increase of the first-order momentum parameter ensures convergence; (ii) it uncovers a phased co-adaptation mechanism between momentum and step size that enables acceleration; and (iii) it extends the analysis to composite, minimax, and bilevel optimization settings. Theoretical contributions include: (i) nonconvex convergence guarantees applicable to broad classes of adaptive methods (e.g., Adam-type); and (ii) novel, efficient minimax and bilevel optimization algorithms that avoid large batch sizes or double-loop schemes. Empirical results confirm improved convergence rates and generalization performance, corroborating the theoretical insights.
📝 Abstract
Although adaptive optimization algorithms have been successful in many applications, there are still some mysteries in terms of convergence analysis that have not been unraveled. This paper provides a novel non-convex analysis of adaptive optimization to uncover some of these mysteries. Our contributions are three-fold. First, we show that an increasing or large enough momentum parameter for the first-order moment used in practice is sufficient to ensure the convergence of adaptive algorithms whose adaptive scaling factors of the step size are bounded. Second, our analysis gives insights for practical implementations, e.g., increasing the momentum parameter in a stage-wise manner in accordance with stagewise decreasing step size would help improve the convergence. Third, the modular nature of our analysis allows its extension to solving other optimization problems, e.g., compositional, min-max and bilevel problems. As an interesting yet non-trivial use case, we present algorithms for solving non-convex min-max optimization and bilevel optimization that do not require using large batches of data to estimate gradients or double loops as the literature do. Our empirical studies corroborate our theoretical results.
Problem

Research questions and friction points this paper is trying to address.

Analyzes convergence of adaptive optimization algorithms
Provides insights for practical implementation improvements
Extends analysis to solve complex optimization problems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Non-convex analysis for adaptive optimization convergence
Stage-wise momentum increase with step size decrease
Modular analysis extends to min-max and bilevel problems
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Texas A&M University | Dalian University of Technology | Alibaba Group
Zhishuai Guo
Zhishuai Guo
Assistant Professor, Northern Illinois University
Federated LearningMachine LearningOptimization
Y
Yi Xu
Dalian University of Technology, Dalian, 116024, Liaoning, China
W
W. Yin
Alibaba Group, Bellevue, WA 98004, USA
R
Rong Jin
Alibaba Group, Bellevue, WA 98004, USA
Tianbao Yang
Tianbao Yang
Texas A&M University
machine learningstochastic optimization