Convergence of Practical Muon with Finite Newton-Schulz Iterations and Nesterov Momentum

๐Ÿ“… 2026-09-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the theoretical gap concerning the non-convex convergence of the practical Muon optimizer under finite Newtonโ€“Schulz iterations with Nesterov momentum. For layer-wise finite-step updates, we propose a coupled learning rate and momentum scheduling strategy, conducting theoretical analysis via descent inequalities and momentum tracking error decomposition techniques. This work establishes the first convergence guarantee for finite-step Muon while preserving the Nesterov recursion. Notably, our analysis eliminates the need for global orthogonalization assumptions, yields convergence constants free of dimensional factors, and relaxes noise conditions. Furthermore, we prove an $O(T^{-1/4})$ convergence rate in terms of the average gradient norm and validate the scalar mapping bounds for the five-step quintic polynomial approximation.
๐Ÿ“ Abstract
Practical Muon maintains momentum and performs a small, fixed number of Newton--Schulz iterations separately for each parameter matrix, often with a Nesterov correction. We analyze these layer-wise finite-step updates jointly on a coupled nonconvex objective, rather than replacing them by exact polar factors or one global orthogonalization. Under gradient-dependent $(\mathcal L_0,\mathcal L_1,q)$-smoothness and conditionally unbiased stochastic gradients with bounded layer-wise variance, we establish an $\mathcal O(T^{-1/4})$ bound on the expected average Frobenius gradient norm. The analysis retains the Nesterov recursion and requires neither bounded stochastic gradients, symmetric noise, nor a uniform positive lower bound on the nonzero output singular values. Its constants contain no explicit matrix-dimension or rank factors when the number of blocks and problem constants are fixed. The proof follows a descent inequality and a decomposition of the momentum tracking error into initialization, noise, and drift. For the original five-step quintic, we verify the required scalar-map bounds analytically; the result also allows step-dependent coefficients satisfying the same bounds. A complementary nuclear-norm result quantifies rank dependence under a stronger spectral condition. The vanishing rate uses coupled learning-rate and momentum schedules, including the standard single-coefficient Nesterov rule.
Problem

Research questions and friction points this paper is trying to address.

Practical Muon
Newton-Schulz iterations
Nesterov momentum
nonconvex optimization
convergence analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Practical Muon
Newton-Schulz iterations
Nesterov momentum
nonconvex optimization
convergence analysis