arXiv · 2609.35701
MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining
Abstract
The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining performance. However, row-wise normalization alone cannot accommodate different imbalance patterns in update matrices. In this paper, we propose an improved Muon optimizer, called \underline{m}atrix-\underline{eq}uilibrating Muon~(MeqMuon), for LLM pretraining. MeqMuon balances both row and column magnitudes through normalization that can be automatically tailored to different imbalance patterns without manual intervention. Moreover, MeqMuon eliminates the need to store AdamW's second-moment estimates, reducing optimizer-state memory usage. Empirical results demonstrate that MeqMuon achieves better convergence performance than AdamW, Muon, and other baselines in LLM pretraining.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chang-Wei Shi, Xu Wang, Wu-Jun Li. 2026-09-28. MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining. https://arxiv.org/abs/2609.35701
Cite the original work for its findings. Save a collection to share your selection of sources.