🤖 AI Summary
This work addresses the limitation of conventional optimizers (e.g., Adam) that rely on hand-crafted gradient estimation heuristics. We propose Trainable Optimizer (TO), a framework that jointly trains a parameterized gradient estimator alongside the model parameters. Crucially, TO incorporates a pseudo-linear approximation of the estimator, enabling SGD-like convergence rates while substantially reducing gradient estimation variance. To enhance computational efficiency, we further introduce two lightweight variants requiring only minimal additional tensor operations. Theoretical analysis establishes convergence guarantees for both strongly convex and non-convex objectives. Empirical evaluation demonstrates that TO achieves faster convergence than Adam and other baselines across diverse benchmark tasks. Moreover, TO exhibits strong efficacy and scalability in fine-tuning large language models, validating its practical utility in modern deep learning settings.
📝 Abstract
The concept of learning to optimize involves utilizing a trainable optimization strategy rather than relying on manually defined full gradient estimations such as ADAM. We present a framework that jointly trains the full gradient estimator and the trainable weights of the model. Specifically, we prove that pseudo-linear TO (Trainable Optimizer), a linear approximation of the full gradient, matches SGD's convergence rate while effectively reducing variance. Pseudo-linear TO incurs negligible computational overhead, requiring only minimal additional tensor multiplications. To further improve computational efficiency, we introduce two simplified variants of Pseudo-linear TO. Experiments demonstrate that TO methods converge faster than benchmark algorithms (e.g., ADAM) in both strongly convex and non-convex settings, and fine tuning of an LLM.