SOLAR: A State-Driven Online Learning Rate Scheduler for LLM Pretraining

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the rigidity of static learning rate schedules in LLM pretraining and the instability inherent in large-scale online learning by proposing the SOLAR framework. The method achieves stable online learning rate adaptation through state-driven bounded residual correction, eliminating the need to retrain warmup-decay curves. It innovatively integrates benchmark anchoring, group-level control, and a circuit breaker mechanism to ensure robustness. Furthermore, SOLAR employs a lightweight PPO strategy guided by progress-aware rewards, maintaining compatibility with AdamW and Muon optimizers as well as Mixture-of-Experts architectures. Evaluated on models ranging from 60M to 3B parameters, the framework significantly reduces perplexity, demonstrating effective policy transferability across scales.
📝 Abstract
Learning-rate (LR) scheduling plays a central role in large language model (LLM) pretraining, yet current practice still relies heavily on hand-crafted heuristics such as Warmup-Cosine-Decay and Warmup-Stable-Decay. Because these schedules are fixed in advance, they cannot adapt to evolving optimization dynamics. Online learned scheduling within the Learning to Optimize (L2O) framework offers a dynamic alternative, but remains brittle at LLM scale due to noisy signals, delayed feedback, and the risk of catastrophic divergence. We propose SOLAR (State-driven Online Learning rAte scheduleR), a stabilized framework for reliable online LR adaptation. SOLAR uses a base schedule as a reference and learns bounded, state-dependent residual corrections for individual parameter groups. Each correction re-anchors to the base at every step, allowing the policy to adapt the LR without relearning the warmup-decay profile. A lightweight state representation and progress-aware reward guide online learning, while a Circuit-Breaker restores training after rare unsafe actions. Across autoregressive language-model pretraining, SOLAR improves final perplexity over tuned static schedules and automatic LR tuners for dense models from 60M to 1B, AdamW and Muon, and two MoE settings up to 3B. Matched 130M controls show that adding base anchoring and action bounds improves a global PPO controller from 27.09 to 23.74 final PPL, while group-wise control reaches 22.87 on the same two seeds. A residual policy trained on a 60M proxy can also be frozen and reused at larger dense scales without target PPO updates, remaining effective across a fourfold base-LR range. These results establish SOLAR as a practical learned LR controller for LLM pretraining.
Problem

Research questions and friction points this paper is trying to address.

Learning Rate Scheduling
LLM Pretraining
Online Learning
Learning to Optimize
Optimization Dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Online Learning Rate Scheduling
Learning to Optimize
Residual Policy
Circuit-Breaker
LLM Pretraining
🔎 Similar Papers
No similar papers found.
Q
Qiulin Shang
Peking University
B
Binyu Wang
Nanjing University
Y
Yongqi Qiao
Peking University
S
Songde Rao
Peking University
Z
Zhoutong Wu
Peking University
Kun Yuan
Kun Yuan
Center for Machine Learning Research, Peking University
distributed signal processinglarge-scale optimizationmachine learning