🤖 AI Summary
Existing optimizer benchmarks are often confined to a single batch size, overlooking the reliability of hyperparameter scaling rules across varying batch sizes. This study systematically investigates the dependence of optimizer performance on batch size through large-scale language model pretraining experiments, comparative evaluations of multiple optimizers, and extensive hyperparameter tuning. It provides the first empirical evidence that no universal Muon scaling rule generalizes across training configurations, and that the optimal optimizer shifts dynamically with batch size. By exposing the failure of prevailing scaling assumptions, this work demonstrates that optimizers must be independently selected and tuned for specific batch sizes to effectively enhance training efficiency.
📝 Abstract
A plethora of new adaptive optimizers are designed to efficiently estimate and use minibatch gradient statistics to shape parameter updates, but they are typically benchmarked at a single batch size. Hyperparameter scaling rules promise to preserve performance as batch size and gradient noise change, suggesting that the best optimizer at one batch size should remain the best at another. We challenge this approach to developing and evaluating optimizers by showing: (1) no principled scaling rule for Muon works consistently across training settings, and (2) the best optimizer for language model pretraining changes with batch size even after extensive hyperparameter tuning.