arXiv · 2609.33252
CascadeEP: Asynchronous Expert Execution for MoE Prefill under Attention Imbalance
Abstract
Mixture-of-experts (MoE) serving commonly deploys data and expert parallelism (DEP): attention replicas run distinct request batches while routed experts are sharded across an expert-parallel (EP) group. During prefill, attention replicas finish dispatch at different times, but synchronous EP delays expert feed-forward network (FFN) computation until routed inputs from all replicas are ready. Request schedulers seek to balance load while reusing the key-value (KV) cache of shared prompt prefixes to avoid redundant prefill computation. These goals can conflict when a replica holding a matching prefix is already overloaded, leaving residual attention imbalance. We present ASYNCEP, a distributed execution engine for MoE prefill. ASYNCEP proposes three mechanisms. Asynchronous EP allows expert computation to start before tokens from all attention replicas are ready. streamFFN batches ready tokens to balance early execution with FFN computation efficiency. Opportunistic expert weight fetching (OEWF) allows a faster replica to fetch expert weights and execute unstarted work from other GPUs. We evaluate ASYNCEP on DeepSeek-V4-Flash, DeepSeek-V4-Pro, and GLM-5.3, and our results show that ASYNCEP achieves up to 1.48x speedup in p95 time-to-first-token (TTFT) and improves the inference throughput by up to 1.17x.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jin Qin, Tiancheng Hu, Shiyan Wang, Junhao Hu, Zexin Jian, Yuzheng Wang, Haoyu Li, Chunwei Xia, Ying Liu, Pixian Zhan, Di Wang, Zhongzhe Hu, Huimin Cui, Tao Xie, Chenxi Wang. 2026-09-27. CascadeEP: Asynchronous Expert Execution for MoE Prefill under Attention Imbalance. https://arxiv.org/abs/2609.33252
Cite the original work for its findings. Save a collection to share your selection of sources.