🤖 AI Summary
This study addresses the unclear adaptive mechanism and the orthogonalization paradox caused by geometric misalignment in the NorMuon optimizer. To overcome these limitations, this work proposes the DGA-Muon optimizer, revealing the scaling degradation phenomenon inherent in NorMuon and establishing two key design principles: decoupling adaptivity from orthogonalization and enforcing shape alignment. Specifically, the proposed method computes scaling factors based on raw gradients and aligns the orthogonal structure with matrix shapes, integrating first-order moment estimation, bias correction, adaptive clipping, and row-column differentiated scaling strategies. Furthermore, this paper provides theoretical convergence guarantees for DGA-Muon. Experimental results validate the correctness of the theoretical analysis and demonstrate that the proposed approach significantly outperforms NorMuon in overall performance.
📝 Abstract
While NorMuon has achieved strong empirical performance in large-scale pretraining by enhancing Muon with row-wise adaptive scaling, its underlying adaptive mechanism remains poorly understood. In this work, we provide the first systematic analysis of NorMuon's adaptivity, revealing that it originates primarily from orthogonalization-induced geometry rather than genuine optimization-relevant information. Under exact orthogonalization, the adaptive scaling factors degenerate into a single global scalar for square and wide matrices, while for tall matrices their variation arises from the non-uniform distribution of row energy after orthogonalization. Under approximate orthogonalization, the orthogonality residual introduces additional variation into the scaling factors, leading to the \textit{Orthogonalization--Adaptivity Paradox}: more accurate orthogonalization weakens adaptivity. We further show that NorMuon's rigid row-wise scaling is geometrically misaligned with the one-sided orthogonal structure of tall matrices. Based on the analysis of these limitations, we propose two core design principles that a desirable adaptive mechanism for Muon should satisfy. First, adaptive scaling should be decoupled from orthogonalization, with the scaling factors computed directly from raw gradients. Second, adaptive scaling should be aligned with the shape-dependent orthogonal structure of the polar factor, using row-wise scaling for wide matrices and column-wise scaling for tall matrices. By incorporating several other design considerations, including sum-based second-moment estimates, bias correction, and adaptive clipping of scaling factors, we obtain the Decoupled Geometry-Aligned Muon (DGA-Muon) optimizer. We establish convergence guarantees for DGA-Muon and empirically validate both our theoretical characterization of NorMuon's scaling degeneration and the superiority of DGA-Muon over NorMuon.