🤖 AI Summary
This study addresses the lack of convergence guarantees for training loss in the Muon optimizer under finite-step Newton-Schulz orthogonalization. Focusing on two-layer ReLU networks, this work integrates Neural Tangent Kernel theory with high-dimensional probability to provide the first analysis combining momentum accumulation with finite-step update tuning. It reveals the interplay between momentum accumulation and gradient-update alignment, establishing precision-independent bounds on sufficient width and hitting time without requiring exact orthogonalization. Theoretically, it proves that Muon can reach arbitrary target losses within finite time. Empirically, experiments demonstrate that its gradient-update alignment surpasses reference baselines, with all runs achieving successful convergence while preserving kernel positive definiteness.
📝 Abstract
Existing convergence analyses of Muon either assume exact orthogonalization or analyze classical Newton--Schulz polynomials, and guarantee only stationarity, so it is unresolved what Muon's five tuned Newton--Schulz steps preserve and whether that suffices to reach a prescribed neural-network training loss. We establish a finite-time training guarantee that accounts for both momentum accumulation before orthogonalization and the tuned finite-step update. For full-batch training of a sufficiently wide two-layer ReLU network with fixed random output weights and a positive-definite limiting neural tangent kernel, we prove that Muon reaches any target empirical squared loss $\varepsilon>0$ with high probability over initialization. For every momentum parameter $μ\in[0,1)$, a target-dependent constant learning rate proportional to $(1-μ)\sqrt{\varepsilon}$ yields a hitting-time bound of $O((1-μ)^{-1}\varepsilon^{-1/2})$, with other problem parameters fixed. The sufficient width is independent of both target accuracy and momentum. The analysis shows that the tuned Newton--Schulz map preserves alignment with the momentum buffer while bounding the update's spectral norm. Control of gradient variation near initialization transfers this alignment to the current gradient, ensuring descent until the target is reached without requiring exact orthogonalization. Numerical experiments support these mechanisms at widths below the sufficient theoretical threshold: gradient-update alignment remains above the analytical reference, and all 30 runs across six widths and five student initializations on a fixed teacher-student dataset reach the target loss while maintaining kernel positivity.