🤖 AI Summary
This study systematically compares the efficiency and effectiveness of AdamW, Lion, and Sophia optimizers for large language model (LLM) pretraining under constrained computational budgets. Methodologically, it employs Maximal Update Parametrization (μP) and a proxy-model paradigm, coupled with single- and multi-cycle learning rate schedules, to enable scalable hyperparameter tuning. Results reveal distinct trade-offs: Sophia achieves the lowest training and validation loss, yielding superior convergence quality; Lion attains the highest training throughput and fastest wall-clock convergence; while AdamW demonstrates strongest generalization on downstream benchmarks (e.g., MMLU, ARC). Crucially, this work establishes the first unified experimental framework to characterize the interplay among optimization dynamics, resource consumption, and final model performance across these three optimizers. The findings provide a reproducible, empirically grounded guideline for optimizer selection in LLM pretraining—bridging theoretical design principles with practical implementation constraints.
📝 Abstract
Optimizers play a decisive role in reducing pre-training times for LLMs and achieving better-performing models. In this study, we compare three major variants: the de-facto standard AdamW, the simpler Lion, developed through an evolutionary search, and the second-order optimizer Sophia. For better generalization, we train with two different base architectures and use a single- and a multiple-epoch approach while keeping the number of tokens constant. Using the Maximal Update Parametrization and smaller proxy models, we tune relevant hyperparameters separately for each combination of base architecture and optimizer. We found that while the results from all three optimizers were in approximately the same range, Sophia exhibited the lowest training and validation loss, Lion was fastest in terms of training GPU hours but AdamW led to the best downstream evaluation results.