🤖 AI Summary
This work addresses the lack of theoretical convergence guarantees in existing adaptive learning rate optimizers, such as Adam, which can compromise training stability. The authors propose C-Adam, a novel optimizer that integrates the geometric properties of line-of-sight methods into an adaptive optimization framework. To the best of our knowledge, this is the first approach to formally incorporate line-of-sight principles into adaptive optimization, and the paper provides a rigorous proof of its convergence. By synergistically combining adaptive learning rates with the directional stability of line-of-sight updates, C-Adam maintains computational efficiency while ensuring theoretical convergence. Empirical evaluations across multiple real-world scenarios demonstrate that C-Adam achieves consistently stable and superior performance, effectively bridging the gap between theoretical soundness and practical efficacy.
📝 Abstract
A crucial component of machine learning algorithms is minimizing loss functions with less computational cost and less oscillations. While adaptive learning rate-based optimizers have been widely used for real-world tasks, they do not guarantee convergence, which is why AMSGrad was later introduced to investigate the non-convergence behaviour of Adam. In this paper, popular adaptive optimization methods like Adam and AMSGrad are critically reviewed with an emphasis on their fundamental design concepts. To address limitations of the above mentioned optimizers, a new optimizer variant, C-Adam, is proposed based on the line of sight approach. A theoretical proof for convergence is also provided and the optimizer is validated through a number of real-life based numerical experiments.