🤖 AI Summary
This study addresses the limitation of the standard Muon optimizer, which lacks a principled treatment of embedding layers and output heads. To resolve this, we propose MuonIO, which unifies and extends the Muon update rule to input and output layers via column- and row-wise normalization. This formulation is grounded in spectral norm theory—specifically the 2→∞ and 1→2 operator norms—and supported by RMS stability arguments alongside local linearization analysis. In 1B-parameter LLaMA pretraining experiments, MuonIO reduces optimizer state memory by 50% and update FLOPs by 46% while simultaneously improving validation perplexity. Overall, this work achieves efficient and unified optimization for both I/O layers, offering a theoretically motivated and practically scalable enhancement to the Muon framework.
📝 Abstract
The Muon optimizer derives its update rule for hidden linear layers by solving a local linearization of the loss penalized by the spectral norm, motivated by an RMS-stability argument for dense linear layers. Standard Muon implementations, however, exclude the input (embedding table) and output (language model head) layers from this principled treatment, for which they use AdamW instead. We present MuonIO, a single Muon-style update for both of these layers. For the language model head $\mathbf{L} \in \mathbb{R}^{V \times d}$, we motivate the use of the $2\to\infty$ operator norm, due to the Lipschitz continuity of the softmax output geometry, while for the embedding table $\mathbf{E} \in \mathbb{R}^{d \times V}$, we draw on the $1 \to 2$ operator norm, based on the one-hot input geometry identified by Bernstein & Newhouse (2025). The identity $\lVert\mathbf{L}\rVert_{2\to\infty}=\lVert\mathbf{L}^\top\rVert_{1\to2}$ then puts both matrices in the same vocabulary-oriented geometry: MuonIO applies a single normalized-vector rule, which appears as column normalization for $\mathbf{E}$ and row normalization for $\mathbf{L}$. Empirical evaluations demonstrate the effectiveness of our approach, with MuonIO reducing I/O optimizer state memory by 50% and I/O update FLOPs by $\sim$46% compared to Muon for 1B LLaMA pretraining on C4, while also improving validation perplexity.