arXiv · 2609.34525
SAGE: Subspace Alignment for Classifier-Free Guidance in Mixture-of-Experts Diffusion Models
Abstract
Diffusion Transformers with Mixture-of-Experts (MoE) routing are a leading recipe for scaling generative models. Classifier-Free Guidance (CFG) is essential for generation quality, yet excessively high guidance scales trigger collapse. We identify a previously unreported failure mode in their combination: the two CFG branches route independently, so their realized activations occupy different subspaces. The unconditional write then leaves the conditional subspace, and CFG amplifies that residual linearly in the guidance scale. We propose SAGE, a training-time regularizer that aligns unconditional MoE activations to the conditional subspace without restricting routing diversity, at zero inference cost. Toy experiments show that SAGE dramatically suppresses extreme drift by 9.2x. When scaled to a 1B-parameter text-to-image model, SAGE significantly improves generation quality, delivering a 9.3% boost in peak DPG-Bench performance. Extensive experiments demonstrate that SAGE consistently outperforms the baseline.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Boyu Zhang, Yangming Cheng, Ning Zhang, Pengfei Liu, Weijie Li, Yifan Gao, Hangyu Li, Litong Gong. 2026-09-28. SAGE: Subspace Alignment for Classifier-Free Guidance in Mixture-of-Experts Diffusion Models. https://arxiv.org/abs/2609.34525
Cite the original work for its findings. Save a collection to share your selection of sources.