CascadeEP: Asynchronous Expert Execution for MoE Prefill under Attention Imbalance

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inference bottleneck in Mixture-of-Experts (MoE) models caused by uneven attention workloads during the prefill phase, which forces synchronous waiting under expert parallelism. To overcome this, we propose ASYNCEP, a distributed engine that introduces an asynchronous expert parallelism mechanism to eliminate global synchronization barriers. The system further enhances computational efficiency through streaming FFN batching and achieves dynamic load balancing via an opportunistic expert weight fetching (OEWF) strategy. By integrating data and expert parallelism with KV cache reuse, ASYNCEP optimizes the end-to-end serving pipeline. Extensive evaluations on DeepSeek and GLM model series demonstrate that our approach reduces p95 time-to-first-token latency by 1.48× and improves throughput by 1.17× compared to existing baselines.
📝 Abstract
Mixture-of-experts (MoE) serving commonly deploys data and expert parallelism (DEP): attention replicas run distinct request batches while routed experts are sharded across an expert-parallel (EP) group. During prefill, attention replicas finish dispatch at different times, but synchronous EP delays expert feed-forward network (FFN) computation until routed inputs from all replicas are ready. Request schedulers seek to balance load while reusing the key-value (KV) cache of shared prompt prefixes to avoid redundant prefill computation. These goals can conflict when a replica holding a matching prefix is already overloaded, leaving residual attention imbalance. We present ASYNCEP, a distributed execution engine for MoE prefill. ASYNCEP proposes three mechanisms. Asynchronous EP allows expert computation to start before tokens from all attention replicas are ready. streamFFN batches ready tokens to balance early execution with FFN computation efficiency. Opportunistic expert weight fetching (OEWF) allows a faster replica to fetch expert weights and execute unstarted work from other GPUs. We evaluate ASYNCEP on DeepSeek-V4-Flash, DeepSeek-V4-Pro, and GLM-5.3, and our results show that ASYNCEP achieves up to 1.48x speedup in p95 time-to-first-token (TTFT) and improves the inference throughput by up to 1.17x.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Prefill
Attention Imbalance
Synchronous Expert Parallelism
Inference Serving
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Asynchronous Expert Parallelism
MoE Prefill
Attention Imbalance
Distributed Inference
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jin Qin
University of the Chinese Academy of Sciences
Tiancheng Hu
Tiancheng Hu
University of Cambridge
natural language processingcomputational social science
S
Shiyan Wang
Beijing University of Posts and Telecommunications
Junhao Hu
Junhao Hu
Peking University
LLM systemsLLM applications
Z
Zexin Jian
University of the Chinese Academy of Sciences
Yuzheng Wang
Yuzheng Wang
Fudan University
Knowledge DistillationVision-Language ModelAIGC
H
Haoyu Li
University of the Chinese Academy of Sciences
C
Chunwei Xia
University of Leeds
Y
Ying Liu
Institute of Computing Technology, Chinese Academy of Sciences
P
Pixian Zhan
Advanced Institute of Information Technology
D
Di Wang
Peking University
Z
Zhongzhe Hu
Huawei Technologies Ltd.
H
Huimin Cui
Institute of Computing Technology, Chinese Academy of Sciences
T
Tao Xie
Beijing Tongming Lake Information Technology Application Innovation Center
Chenxi Wang
Chenxi Wang
Institute of Computing Technology, Chinese Academy of Sciences
Operating SystemManaged RuntimeProgramming Language