arXiv · 2610.06116
ORCA: The Annealed Spectral Conditioning Optimizer for Faster, Better LLM Training
Abstract
Modern LLM optimizers such as Muon often produce weight matrices with higher effective rank than Adam, yet further spectral control has delivered only modest gains. We identify a tension behind this result: concentrated spectra can suppress gradient directions in coupled weight matrices and slow optimization, while constraints maintained throughout training can limit task-specific adaptation and raise the attainable loss floor. We introduce ORCA (Orthogonal Regularization, Cooled After), a minimal optimizer intervention that applies strong but temporary soft orthogonality regularization early in training, then removes it. This allows the weights to benefit from a broader spectrum early on and adapt freely afterward. Across LLaMA, Qwen3, and fine-grained mixture-of-experts models ranging from 130M to 8B parameters, ORCA achieves lower final validation loss than Muon. Its loss reduction relative to Muon matches or exceeds Muon's reduction relative to Adam. Ablations support the early-shaping, later-release design. Further, ORCA requires no architectural changes and adds minimal overhead.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuanshi Liu, Boyuan Jiang, Liang Hou, Xin Tao, Pengfei Wan, Zhouchen Lin, Cong Fang. 2026-10-05. ORCA: The Annealed Spectral Conditioning Optimizer for Faster, Better LLM Training. https://arxiv.org/abs/2610.06116
Cite the original work for its findings. Save a collection to share your selection of sources.