🤖 AI Summary
This study addresses the longstanding reliance on heuristic settings for learning rate warmup duration in large model training, which lacks theoretical justification for dynamic scaling with training cycles. By modeling quadratic loss landscapes and conducting stability margin analysis, this work reveals the trade-off mechanism between early progress sacrifice and long-term gains. It derives a period-dependent scaling law for warmup duration, transforming it from a fixed empirical value into a hyperparameter contingent upon the training cycle. This theoretical framework unifies the explanation of diverse warmup phenomena and enables accurate prediction of optimal warmup durations for extended training cycles using only short-horizon experiments. Consequently, the proposed approach significantly enhances the training efficiency of large-scale models.
📝 Abstract
Learning-rate warmup is a standard technique in language-model training, yet its duration remains largely heuristic. Common approaches use either a fixed number of updates or a fixed fraction of the training horizon, two choices that imply very different scaling as training gets longer. When should warmup stay fixed, and when should it grow with the horizon? We address this question with a quadratic model whose modes respond differently to the peak learning rate. Warmup slows progress in directions that already contract well at the peak rate, but can remove persistent error in directions near the stability edge, with higher peak rates shifting the balance toward longer warmup durations. This yields a compact horizon scaling law that captures regimes ranging from essentially no warmup, through fixed-duration warmup, to durations that grow with the training horizon, and explains how the preferred regime changes with peak learning rate. Because the law captures the tradeoff between giving up early progress and improving the trajectory that follows, it can be fit using shorter runs and used to predict warmup at substantially longer horizons. Together, our results explain several familiar properties of warmup through a single tradeoff and suggest treating warmup duration as a horizon-dependent hyperparameter rather than a fixed training heuristic.