🤖 AI Summary
This study addresses the coupling of marginal scale and interaction geometry in matrix optimizers for large language model training. To overcome this limitation, it introduces a novel Normalize-Then-Precondition paradigm and proposes the NormPre-G/L hierarchical optimization framework. This approach decouples update directions via diagonal Gram normalization and combines Newton-Schulz iterations with randomized sketching to approximate the principal eigenspace, thereby achieving efficient second-order optimization through spectral norm steepest descent. The convergence of the proposed algorithm is theoretically established. Empirical evaluations on pretraining tasks, including GPT-2, demonstrate that NormPre-G/L outperforms AdamW and Muon. The source code has been made publicly available.
📝 Abstract
Matrix optimizers have emerged as a promising direction, with Muon standing out as a prominent design. Revisiting Muon through its full-Gram representation, we observe that it jointly processes marginal-scale and interaction information. This opens an alternative way to organize geometric information hierarchically, motivating the Normalize-Then-Precondition framework. Specifically, it first uses diagonal-Gram information to construct a marginally normalized update, then applies spectral preconditioning to its directional interaction geometry. Building on this framework, we develop NormPre with NormPre-G and NormPre-L adopting global and localized spectral preconditioning, grounded in spectral-norm steepest descent and a regularized formulation followed by leading mode selection, respectively. To enable large-scale training, NormPre-G uses Newton-Schulz iterations and NormPre-L employs randomized sketching to approximate the leading interaction eigenspace. Theoretically, we establish $\mathcal{O}(T^{-1/2})$ convergence guarantees for simplified versions of NormPre. Across extensive pretraining experiments on GPT-2 Small, LLaMA and Qwen3, both variants consistently outperform AdamW, Muon and MANO under matched training budgets. Further efficiency and spectral analyses reveal the complementary strengths of two variants and characterize their performance-efficiency trade-off. We open-source our code through a GitHub repository at https://github.com/zx-gong/NormPre.