🤖 AI Summary
This study addresses the limitations of existing optimizers in adaptively balancing row and column scales of update matrices and their high memory consumption. To this end, we propose MeqMuon, an optimizer that introduces adaptive row-column normalization to automatically equilibrate the magnitudes across rows and columns of the update matrix without manual hyperparameter tuning. Furthermore, MeqMuon eliminates the storage of second-order moments, substantially reducing the memory overhead associated with optimizer states. Experimental results demonstrate that MeqMuon achieves superior convergence performance compared to mainstream baselines, including AdamW and Muon, on large language model pretraining tasks. By effectively reconciling optimization efficacy with minimal memory footprint, the proposed method enables highly efficient training under stringent memory constraints.
📝 Abstract
The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining performance. However, row-wise normalization alone cannot accommodate different imbalance patterns in update matrices. In this paper, we propose an improved Muon optimizer, called \underline{m}atrix-\underline{eq}uilibrating Muon~(MeqMuon), for LLM pretraining. MeqMuon balances both row and column magnitudes through normalization that can be automatically tailored to different imbalance patterns without manual intervention. Moreover, MeqMuon eliminates the need to store AdamW's second-moment estimates, reducing optimizer-state memory usage. Empirical results demonstrate that MeqMuon achieves better convergence performance than AdamW, Muon, and other baselines in LLM pretraining.