🤖 AI Summary
This study addresses the inference bottleneck in Mixture-of-Experts (MoE) models caused by uneven attention workloads during the prefill phase, which forces synchronous waiting under expert parallelism. To overcome this, we propose ASYNCEP, a distributed engine that introduces an asynchronous expert parallelism mechanism to eliminate global synchronization barriers. The system further enhances computational efficiency through streaming FFN batching and achieves dynamic load balancing via an opportunistic expert weight fetching (OEWF) strategy. By integrating data and expert parallelism with KV cache reuse, ASYNCEP optimizes the end-to-end serving pipeline. Extensive evaluations on DeepSeek and GLM model series demonstrate that our approach reduces p95 time-to-first-token latency by 1.48× and improves throughput by 1.17× compared to existing baselines.
📝 Abstract
Mixture-of-experts (MoE) serving commonly deploys data and expert parallelism (DEP): attention replicas run distinct request batches while routed experts are sharded across an expert-parallel (EP) group. During prefill, attention replicas finish dispatch at different times, but synchronous EP delays expert feed-forward network (FFN) computation until routed inputs from all replicas are ready. Request schedulers seek to balance load while reusing the key-value (KV) cache of shared prompt prefixes to avoid redundant prefill computation. These goals can conflict when a replica holding a matching prefix is already overloaded, leaving residual attention imbalance. We present ASYNCEP, a distributed execution engine for MoE prefill. ASYNCEP proposes three mechanisms. Asynchronous EP allows expert computation to start before tokens from all attention replicas are ready. streamFFN batches ready tokens to balance early execution with FFN computation efficiency. Opportunistic expert weight fetching (OEWF) allows a faster replica to fetch expert weights and execute unstarted work from other GPUs. We evaluate ASYNCEP on DeepSeek-V4-Flash, DeepSeek-V4-Pro, and GLM-5.3, and our results show that ASYNCEP achieves up to 1.48x speedup in p95 time-to-first-token (TTFT) and improves the inference throughput by up to 1.17x.