From Geometry to Generalization: Why Row Normalization Can Beat Adam and Muon

πŸ“… 2026-10-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the unclear mechanism by which geometric discrepancies among optimizers affect generalization, noting that mainstream optimizers such as Adam and Muon are prone to geometric distortion in high-dimensional classification. Leveraging high-dimensional multiclass classification theory and power-law spectral data modeling, this work provides the first rigorous geometric analysis of optimizer generalization behavior under anisotropic data, validated through synthetic experiments and last-layer fine-tuning of language models. It demonstrates that row normalization effectively prevents distortions arising from Adam’s coordinate geometry and Muon’s spectral geometry by preserving decision boundary orientations. Under specific covariance structures, this approach achieves population accuracy superior to mainstream optimizers, revealing the underlying geometric mechanisms through which optimizer selection influences model generalization.
πŸ“ Abstract
Different optimizers can fit the same training data while selecting classifiers with substantially different geometries, but whether this difference provably affects population performance remains unclear. We show that row-wise normalization can achieve strictly higher population accuracy than full-batch Adam, a proxy for random-reshuffling Adam, and exact-SVD Muon in high-dimensional multiclass classification. Under an isotropic Gaussian-cloud data model, this advantage arises because row normalization's class-wise Euclidean geometry asymptotically preserves the population decision-boundary directions, whereas Adam's coordinate-wise geometry and Muon's spectral geometry introduce nonvanishing distortions. Beyond isotropy, the advantage persists for full-batch training on class means with independently oriented class-mean and test-noise covariances. It holds for power-law spectra with class-mean exponent below one, even under heavily anisotropic test noise. When both covariances are diagonal and sufficiently close, the advantage over Adam can reverse, while applying the same random rotation to both restores it by changing only their alignment with Adam's coordinate axes. Synthetic and last-layer language-model experiments support the predicted advantage.
Problem

Research questions and friction points this paper is trying to address.

optimizers
generalization
geometry
multiclass classification
population accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Row Normalization
Optimizer Geometry
Generalization
Adam
Muon