Optimizing Large Language Models with Chained LMOs

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of a unified theoretical framework for existing chained matrix normalization optimizers and their risk of divergence on smooth convex objectives. We propose a chained Linear Minimization Oracle (LMO) framework, leveraging linear associative memory to elucidate its effectiveness mechanism and advantages under anisotropic embeddings. Furthermore, we introduce the first cross-layer stacked weight 3D tensor axis normalization method, yielding the TensorChain optimizer. Evaluated during Qwen3 pretraining, our approach achieves superior average token efficiency compared to all baselines, reducing token consumption by 9.6% relative to Muon at matched validation loss. This work establishes both theoretical unification and empirical performance improvements for large language model training optimization.
📝 Abstract
Muon has motivated a growing family of optimizers that compose multiple matrix normalizations, but these methods remain fragmented and lack a unified perspective. We introduce chained linear minimization oracles (chained LMOs), which cast these methods as compositions of LMOs. Despite their empirical success, many chains fall outside the standard LMO framework and can diverge on smooth convex objectives. To explain why composition can nevertheless help, we turn to linear associative memory and show that chaining can improve over Muon under anisotropic embeddings. Empirically, we propose TensorChain, a novel optimizer within the framework that stacks compatible weight matrices across different layers and normalizes the 3d tensor across its axes. In Qwen3 0.6B and 1.7B pretraining, TensorChain outperforms all chained baselines in average token efficiency, with average token savings of 9.6% over Muon at matched validation loss.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Optimization
Linear Minimization Oracles
Muon
Chained LMOs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chained LMOs
TensorChain
Linear Associative Memory
Anisotropic Embeddings
Large Language Model Optimization
🔎 Similar Papers
No similar papers found.