🤖 AI Summary
This study addresses the rigidity of static learning rate schedules in LLM pretraining and the instability inherent in large-scale online learning by proposing the SOLAR framework. The method achieves stable online learning rate adaptation through state-driven bounded residual correction, eliminating the need to retrain warmup-decay curves. It innovatively integrates benchmark anchoring, group-level control, and a circuit breaker mechanism to ensure robustness. Furthermore, SOLAR employs a lightweight PPO strategy guided by progress-aware rewards, maintaining compatibility with AdamW and Muon optimizers as well as Mixture-of-Experts architectures. Evaluated on models ranging from 60M to 3B parameters, the framework significantly reduces perplexity, demonstrating effective policy transferability across scales.
📝 Abstract
Learning-rate (LR) scheduling plays a central role in large language model (LLM) pretraining, yet current practice still relies heavily on hand-crafted heuristics such as Warmup-Cosine-Decay and Warmup-Stable-Decay. Because these schedules are fixed in advance, they cannot adapt to evolving optimization dynamics. Online learned scheduling within the Learning to Optimize (L2O) framework offers a dynamic alternative, but remains brittle at LLM scale due to noisy signals, delayed feedback, and the risk of catastrophic divergence. We propose SOLAR (State-driven Online Learning rAte scheduleR), a stabilized framework for reliable online LR adaptation. SOLAR uses a base schedule as a reference and learns bounded, state-dependent residual corrections for individual parameter groups. Each correction re-anchors to the base at every step, allowing the policy to adapt the LR without relearning the warmup-decay profile. A lightweight state representation and progress-aware reward guide online learning, while a Circuit-Breaker restores training after rare unsafe actions. Across autoregressive language-model pretraining, SOLAR improves final perplexity over tuned static schedules and automatic LR tuners for dense models from 60M to 1B, AdamW and Muon, and two MoE settings up to 3B. Matched 130M controls show that adding base anchoring and action bounds improves a global PPO controller from 27.09 to 23.74 final PPL, while group-wise control reaches 22.87 on the same two seeds. A residual policy trained on a 60M proxy can also be frozen and reused at larger dense scales without target PPO updates, remaining effective across a fourfold base-LR range. These results establish SOLAR as a practical learned LR controller for LLM pretraining.