Muon Sublates the Edge of Stability in LLM Pretraining

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of classical Edge of Stability (EoS) theory for gradient descent to explain the large-step-size dynamics of the Muon optimizer in large-model pretraining. We propose a novel "split EoS" paradigm. Through stochastic momentum-free Muon analysis and coherence-corrected bound derivation, we provide the first proof that loss balancing and temporal alignment in Muon respond to hyperparameters in a decoupled manner. Controlled experiments on Llama architectures ranging from 130M to 1B parameters confirm the existence of weak negative alignment during training, wherein performance continues improving despite incomplete directional reversal. This work revises conventional stability assumptions and reveals the unique dynamical mechanisms underlying Muon. The implementation code has been made publicly available.
📝 Abstract
Muon is increasingly used for language-model pretraining, yet its large-step dynamics are not captured by the classical edge-of-stability (EoS) picture of gradient descent (GD). In GD, loss neutrality, equal-magnitude update reversal, and marginal stability meet at a single learning-rate-dependent edge. We show that Muon breaks this coupling. For stochastic no-momentum Muon, we derive a coherence-corrected conditional loss-neutral boundary $2\rho_b/\eta$, while temporal alignment follows a separate geometry. Controlled experiments show that loss balance and temporal alignment respond differently to learning rate and batch size. Across our language model experiments, the 130M Llama-like LLM runs exhibit loss-boundary tracking with weak negative alignment, whereas the studied 1B LLM configuration shows stronger partial cancellation; in both settings, directions remain far from coherent reversal while training continues to improve. These results support a split EoS picture for Muon: a stochastic loss-neutral edge survives, but it is not accompanied by a universal temporal-direction signature. The source code for reproducing the experiments can be found in https://github.com/cyzebra/Muon-Sublates-the-Edge-of-Stability-in-LLM-Pretraining
Problem

Research questions and friction points this paper is trying to address.

Muon optimizer
Edge of Stability
LLM pretraining
gradient descent
loss neutrality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Muon optimizer
Edge of Stability
LLM pretraining
loss-neutral boundary
stochastic gradient descent
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yanzhe Chen
School of Mathematical Sciences, Shanghai Jiao Tong University, China
Q
Qifang Zhao
Alibaba Inc., Hangzhou, China
X
Xiaoxiao Xu
Alibaba Inc., Hangzhou, China
Fanghui Liu
Fanghui Liu
Assistant Professor, University of Warwick
Foundations of Modern ML