Convergence of Practical Muon

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the theoretical gap surrounding the practical Muon algorithm in large-scale training by providing its first rigorous convergence analysis. Methodologically, we interpret the practical Muon—incorporating Newton-Schulz iterations and decoupled weight decay—as a right-preconditioned stochastic optimizer with dynamic regularization, thereby establishing non-convex convergence guarantees. Our primary contributions include a convergence proof that accounts for practical polynomial coefficients and decoupled weight decay, yielding an O(T^{-1/4}) convergence rate. Furthermore, we improve the dimension-dependence factor to sqrt(d), outperforming AdamW. Experimental results validate the effectiveness of our theoretical findings.
📝 Abstract
Muon is emerging as a promising alternative to AdamW for large-scale neural network training, yet theoretical understanding of its practical implementation remains incomplete, as existing analyses often simplify or omit two key components: (i) practical Newton--Schulz iterations with empirically tuned polynomial coefficients $(3.4445,-4.7750,2.0315)$; and (ii) decoupled weight decay for regularization. In this paper, we provide an optimization interpretation and establish convergence for practical Muon, jointly accounting for both components. Specifically, we interpret practical Muon as right-preconditioned optimization of the original loss with a dynamic weighted $\ell_2$ regularizer that vanishes as stationarity is approached, so that the optimization target remains the original objective. We then establish, to our best knowledge, the first convergence guarantee for practical Muon in the stochastic nonconvex setting, with an $\mathcal{O}(T^{-1/4})$ convergence rate in terms of the expected Frobenius norm of the gradient, improving the dimension dependence of the best known AdamW's convergence rate by a factor of $\sqrt{d}$, where $T$ is the iteration horizon and $d$ is the parameter dimension. Experiments further support the theoretical convergence results.
Problem

Research questions and friction points this paper is trying to address.

Muon optimizer
convergence guarantee
stochastic nonconvex optimization
Newton-Schulz iterations
decoupled weight decay
Innovation

Methods, ideas, or system contributions that make the work stand out.

Muon optimizer
Newton-Schulz iteration
convergence guarantee
stochastic nonconvex optimization
decoupled weight decay
🔎 Similar Papers
No similar papers found.