π€ AI Summary
This study addresses the challenge that existing LLM optimizers struggle to balance spectral control with task adaptability, where suppressing gradient directions often leads to suboptimal loss floors. To overcome this limitation, we propose ORCA, an optimizer introducing a novel "shape early, release late" annealed spectral conditioning strategy. By integrating soft orthogonal regularization with a time-decay schedule, ORCA broadens the weight spectrum during early training to improve optimization dynamics, and subsequently removes constraints to allow unrestricted task adaptation without architectural modifications. Evaluated across model scales from 130M to 8B parameters, including LLaMA and Qwen3, ORCA achieves significantly lower final losses than Muon. Notably, its performance gains over Muon match or exceed Muonβs improvements over Adam, while maintaining low computational overhead.
π Abstract
Modern LLM optimizers such as Muon often produce weight matrices with higher effective rank than Adam, yet further spectral control has delivered only modest gains. We identify a tension behind this result: concentrated spectra can suppress gradient directions in coupled weight matrices and slow optimization, while constraints maintained throughout training can limit task-specific adaptation and raise the attainable loss floor. We introduce ORCA (Orthogonal Regularization, Cooled After), a minimal optimizer intervention that applies strong but temporary soft orthogonality regularization early in training, then removes it. This allows the weights to benefit from a broader spectrum early on and adapt freely afterward. Across LLaMA, Qwen3, and fine-grained mixture-of-experts models ranging from 130M to 8B parameters, ORCA achieves lower final validation loss than Muon. Its loss reduction relative to Muon matches or exceeds Muon's reduction relative to Adam. Ablations support the early-shaping, later-release design. Further, ORCA requires no architectural changes and adds minimal overhead.