Pre-Training LLMs on a budget: A comparison of three optimizers

📅 2025-07-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically compares the efficiency and effectiveness of AdamW, Lion, and Sophia optimizers for large language model (LLM) pretraining under constrained computational budgets. Methodologically, it employs Maximal Update Parametrization (μP) and a proxy-model paradigm, coupled with single- and multi-cycle learning rate schedules, to enable scalable hyperparameter tuning. Results reveal distinct trade-offs: Sophia achieves the lowest training and validation loss, yielding superior convergence quality; Lion attains the highest training throughput and fastest wall-clock convergence; while AdamW demonstrates strongest generalization on downstream benchmarks (e.g., MMLU, ARC). Crucially, this work establishes the first unified experimental framework to characterize the interplay among optimization dynamics, resource consumption, and final model performance across these three optimizers. The findings provide a reproducible, empirically grounded guideline for optimizer selection in LLM pretraining—bridging theoretical design principles with practical implementation constraints.

Technology Category

Machine Learning: OptimizationSearch and Optimization: Learning to SearchNatural Language Processing: Learning & Optimization for NLP

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Optimizers play a decisive role in reducing pre-training times for LLMs and achieving better-performing models. In this study, we compare three major variants: the de-facto standard AdamW, the simpler Lion, developed through an evolutionary search, and the second-order optimizer Sophia. For better generalization, we train with two different base architectures and use a single- and a multiple-epoch approach while keeping the number of tokens constant. Using the Maximal Update Parametrization and smaller proxy models, we tune relevant hyperparameters separately for each combination of base architecture and optimizer. We found that while the results from all three optimizers were in approximately the same range, Sophia exhibited the lowest training and validation loss, Lion was fastest in terms of training GPU hours but AdamW led to the best downstream evaluation results.
Problem

Research questions and friction points this paper is trying to address.

Compare optimizers for efficient LLM pre-training
Evaluate performance of AdamW, Lion, and Sophia
Assess optimizer impact on training speed and model quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compare AdamW, Lion, Sophia optimizers for LLMs
Use Maximal Update Parametrization for tuning
Train with different architectures and epochs
🔎 Similar Papers
J
Joel Schlotthauer
Fraunhofer Institute for Integrated Circuits IIS, Erlangen, Germany
Christian Kroos
Christian Kroos
Fraunhofer Institute for Integrated Circuits
Machine learningEvolutionary computationHuman-robot interactionAuditory-visual speechSpeech production
C
Chris Hinze
Fraunhofer Institute for Integrated Circuits IIS, Erlangen, Germany
Viktor Hangya
Viktor Hangya
Fraunhofer IIS
Natural Language ProcessingMachine Learning
L
Luzian Hahn
Fraunhofer Institute for Integrated Circuits IIS, Erlangen, Germany
F
Fabian Küch
Fraunhofer Institute for Integrated Circuits IIS, Erlangen, Germany