arXiv · 2601.21349
L2R: Low-Rank and Lipschitz-Controlled Routing for Mixture-of-Experts
Abstract
Mixture-of-Experts (MoE) models scale neural networks by conditionally activating a small subset of experts, where the router plays a central role in determining expert specialization and overall model performance. However, many modern MoE systems still adopt linear routers in raw high-dimensional representation spaces, where representation mismatch, angular concentration, and scale-sensitive scoring can jointly undermine routing discriminability and stable expert specialization. In this work, we propose Low-rank & Lipschitz-controlled Routing (L2R), a unified routing framework that reshapes both the routing space and scoring geometry. L2R performs expert assignment in a shared low-rank latent routing space and introduces Saturated Inner-Product Scoring (SIPS) to explicitly control the Lipschitz behavior of routing functions, yielding smoother and more stable routing geometry. In addition, L2R incorporates a parameter-efficient multi-anchor routing mechanism to enhance expert expressiveness. Extensive experiments on an OLMoE-based language MoE model and a vision MoE setting on ImageNet demonstrate that L2R consistently improves routing geometry, expert discrimination, and overall model performance. Code will be released.
Explore related subjects
Keep this discovery
Minghao Yang, Ren Togo, Guang Li, Takahiro Ogawa, Miki Haseyama. 2026-01-29. L2R: Low-Rank and Lipschitz-Controlled Routing for Mixture-of-Experts. https://arxiv.org/abs/2601.21349
Cite the original work for its findings. Save a collection to share your selection of sources.