🤖 AI Summary
This study investigates the gap between theoretical guarantees and empirical performance of row normalization in the NorMuon optimizer during large language model (LLM) pretraining. Leveraging operator norm geometry and polar decomposition, we establish algorithm-dependent upper and lower complexity bounds under both deterministic and stochastic settings, revealing that row normalization introduces a dimension-dependent factor that degrades the theoretical convergence rate. Our analysis demonstrates that, despite this convergence penalty, row normalization substantially improves practical LLM training performance. These findings offer new insights into understanding the discrepancy between theoretical predictions and experimental observations regarding row normalization in adaptive optimization for foundation models.
📝 Abstract
This paper examines how row-wise renormalization affects Muon, focusing on the gap between NorMuon's worst-case guarantees and its practical performance (Li et al.). Despite its growing adoption and promising performance in large language model (LLM) pretraining, NorMuon's worst-case guarantees remain poorly understood. One fundamental question is: Does row normalization yield provable convergence gains, potentially through its interaction with approximate polar computation and exponential moving-average momentum? Our results show that row normalization introduces a dimension-dependent factor in the worst-case iteration complexity under the operator-norm geometry, which persists even with exact polar computation and any fixed momentum parameters. Indeed, we establish an algorithm-dependent lower bound and a matching upper bound in deterministic settings, and extend our upper bound analysis to stochastic settings. Both upper-bound analyses allow approximate polar computation. Experiments show that NorMuon is slower than Muon on synthetic problems inspired by our worst-case construction, yet outperforms Muon in LLM pretraining. These findings sharpen the puzzle of why row normalization helps in practice and complement the recent findings of Dewulf et al.