A Trainable Optimizer

📅 2025-08-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of conventional optimizers (e.g., Adam) that rely on hand-crafted gradient estimation heuristics. We propose Trainable Optimizer (TO), a framework that jointly trains a parameterized gradient estimator alongside the model parameters. Crucially, TO incorporates a pseudo-linear approximation of the estimator, enabling SGD-like convergence rates while substantially reducing gradient estimation variance. To enhance computational efficiency, we further introduce two lightweight variants requiring only minimal additional tensor operations. Theoretical analysis establishes convergence guarantees for both strongly convex and non-convex objectives. Empirical evaluation demonstrates that TO achieves faster convergence than Adam and other baselines across diverse benchmark tasks. Moreover, TO exhibits strong efficacy and scalability in fine-tuning large language models, validating its practical utility in modern deep learning settings.

Technology Category

Machine Learning: OptimizationSearch and Optimization: Learning to SearchNatural Language Processing: Learning & Optimization for NLP

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
The concept of learning to optimize involves utilizing a trainable optimization strategy rather than relying on manually defined full gradient estimations such as ADAM. We present a framework that jointly trains the full gradient estimator and the trainable weights of the model. Specifically, we prove that pseudo-linear TO (Trainable Optimizer), a linear approximation of the full gradient, matches SGD's convergence rate while effectively reducing variance. Pseudo-linear TO incurs negligible computational overhead, requiring only minimal additional tensor multiplications. To further improve computational efficiency, we introduce two simplified variants of Pseudo-linear TO. Experiments demonstrate that TO methods converge faster than benchmark algorithms (e.g., ADAM) in both strongly convex and non-convex settings, and fine tuning of an LLM.
Problem

Research questions and friction points this paper is trying to address.

Develop trainable optimizer replacing manual gradient methods
Prove pseudo-linear TO matches SGD convergence with less variance
Enhance efficiency via simplified TO variants for faster convergence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Trainable Optimizer replaces manual gradient estimations
Pseudo-linear TO matches SGD with lower variance
Simplified TO variants enhance computational efficiency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ruiqi Wang
IEMS, Northwestern University
Diego Klabjan
Diego Klabjan
Northwestern University
Machine learning