🤖 AI Summary
This study addresses the lack of theoretical guidance for hyperparameter tuning in large language model pretraining. By leveraging geometric analysis of local loss landscapes and gradient noise scale theory, it reveals a two-stage dynamic of landscape evolution during pretraining and elucidates the underlying depth-flatness trade-off mechanism. Building upon these insights, this work proposes learning rate warmup and dynamic batch size scheduling strategies that translate landscape evolution patterns into actionable tuning protocols. Ultimately, this research provides an efficient and theoretically grounded guide for hyperparameter optimization in large-scale pretraining.
📝 Abstract
The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase I, sharpness of the local landscape is initially high, leading to instability and loss plateaus under large learning rates (LRs). The landscape shifts from sharp to flatter regions early in training. This dynamic explains the necessity of LR warmup and further suggests that larger peak LRs require proportionally longer warmup periods. In Phase II, the local landscape is governed by the gradient noise scale. Our theory identifies a depth flatness trade-off: high noise from smaller batches widens the loss basin, whereas reduced noise from larger batches deepens it. This theory motivates a dynamic batch-size (BS) scheduler that begins with a small BS and increases it late in training. Together, we provide a unified view of loss landscape evolution, which translates into actionable tuning strategies for large-scale pre-training.