arXiv · 2609.32861
Muon Under Gradient Noise and the Limits of Orthogonalization Near Optima
Abstract
Muon replaces the momentum buffer of each weight matrix by its orthogonal polar factor. We ask what this orthogonalization does near an optimum, where minibatch noise dominates the gradient. Under a Gaussian noise model, the expected Muon update becomes a scaled gradient step, so a linear method with a suitably matched learning rate reproduces Muon's first-order mean response. The stochastic update is a different matter: after the response is matched, Muon retains a nonlinear residual that is uncorrelated with the input noise and contributes additional covariance. A Hermite expansion shows how momentum acts on this residual. Its higher-order components decorrelate faster than the linear component, so momentum suppresses the residual's accumulated covariance relative to the linear part, but never eliminates it. In a local quadratic surrogate that evaluates the residual on the stationary noise buffer, the residual adds stationary covariance and raises the stationary loss floor at every stable step size, while leaving the contraction dynamics unchanged. Simulations of the full nonlinear recursion on quadratics and measurements on frozen transformer gradients support each step of this picture. Together, the results make a theoretical case for replacing orthogonalization by response-matched momentum SGD once optimization becomes noise-dominated.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xiaohui Xie. 2026-09-26. Muon Under Gradient Noise and the Limits of Orthogonalization Near Optima. https://arxiv.org/abs/2609.32861
Cite the original work for its findings. Save a collection to share your selection of sources.