LionMuon: Alternating Spectral and Sign Descent for Efficient Training

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the substantial computational and communication overhead of the Muon optimizer in large model pre-training, as well as the limited directional information inherent to sign-based optimizers. To overcome these bottlenecks, we propose LionMuon, which introduces a novel alternating mechanism between spectral descent and sign descent. By sharing dual exponential moving average (EMA) momentum buffers across low-frequency Muon spectral steps and high-frequency Lion sign steps, the method significantly reduces state memory overhead. Furthermore, we derive theoretical convergence complexity bounds under heavy-tailed noise. Experiments demonstrate that, under identical token budgets, LionMuon achieves superior loss performance compared to mainstream optimizers. In a four-GPU setting, it accelerates training by 33% over Muon, outperforming variants such as Dion while preserving gradient fidelity.
📝 Abstract
Pretraining a language model takes enormous compute, and the right optimizer can save a good part of it. Muon's spectral step gives a stronger direction than a sign step, but it is expensive. Every step runs Newton-Schulz iterations on the full matrix and, in distributed training, an extra all-reduce. Sign steps, as in Lion and Signum, are cheap and stay local to each device. We propose LionMuon, which takes one Muon step every $P$ iterations and Lion steps in between, with a single dual-EMA momentum buffer shared by both. Muon's compute and communication are paid once per $P$ steps, and the optimizer state is half of AdamW's. A single-EMA variant, SignMuon, already improves on Muon. We prove complexity bounds under heavy-tailed noise in which the period sets an interpolation between Muon's and Lion's smoothness and noise constants, and which say when LionMuon is faster than both. On 124M and 355M models trained on FineWeb, LionMuon with $P=2$ and $P=5$ reaches a lower loss than Muon, AdamW, Lion and Signum at the same number of tokens. Under 4-GPU data-parallel training it reaches Muon's final loss with a third less wall-clock on PCIe, and it beats the communication-efficient Muon variants Dion and MuonBP on loss at no more exposed communication, while keeping the exact gradient. Code: https://github.com/brain-lab-research/lion-muon
Problem

Research questions and friction points this paper is trying to address.

Language Model Pretraining
Optimizer Efficiency
Spectral Descent
Sign Descent
Distributed Training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Alternating Spectral and Sign Descent
Dual-EMA Momentum Buffer
Communication Efficiency
Heavy-tailed Noise Complexity Bounds
Distributed Training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Arman Bolatov
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Artem Riabinin
Artem Riabinin
PhD, KAUST
Optimization
N
Nikita Kornilov
Basic Research of Artificial Intelligence Laboratory (BRAIn Lab)
Andrey Veprikov
Andrey Veprikov
Unknown affiliation
OptimizationMLDL
S
Samuel Horváth
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Martin Takáč
Martin Takáč
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
optimizationmachine learningdeep neural networkbig datacomputer science
Aleksandr Beznosikov
Aleksandr Beznosikov
PhD, Basic Research of Artificial Intelligence Lab
OptimizationMachine Learning