Zepp: Accelerating Distributed MoE Serving under Relaxed Balance Constraints

πŸ“… 2026-10-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the computational, communication, and memory imbalances caused by skewed expert loads in distributed Mixture-of-Experts (MoE) serving. We propose a system framework that formulates load balancing as a hard constraint rather than an optimization objective. The core innovation lies in introducing the concept of relaxed balance constraints, which jointly optimizes expert placement, token routing, and execution scheduling to co-adapt to dynamic workloads. Furthermore, we design Split/Merge primitives alongside communication overlapping mechanisms to enable efficient expert parallelism under GPU and NIC resource constraints. Experimental results demonstrate that the proposed system achieves up to 6.68Γ— speedup over seven state-of-the-art baselines, with a geometric mean speedup of 1.86Γ—.
πŸ“ Abstract
As Mixture-of-Experts (MoE) models continue to scale, serving them increasingly relies on expert parallelism (EP) across a growing number of devices. Yet skewed expert workloads create imbalance across computation, communication, and memory, making load balancing a central optimization objective in distributed MoE serving. We observe that balance is not free: operations introduced to balance one dimension can themselves be expensive or imbalanced. This motivates us to rethink balance as a constraint rather than an optimization objective. We present Zepp, which directly optimizes the bottleneck communication in distributed MoE serving subject to simplified balance constraints on physical resources, i.e., GPUs and NICs. Zepp progressively optimizes inter-node communication across placement, routing, and execution. It first places expert replicas to reduce token communication under GPU constraints, then reshapes communication flows through split and merge primitives under NIC constraints, and finally partitions and schedules ex- pert computation to overlap the resulting communication. To adapt to dynamic workloads, Zepp jointly coordinates computation, token communication, and expert-weight movement at each iteration. Together, these designs allow Zepp to pursue the most efficient execution rather than a single-dimension balanced one. We implement Zepp and evaluate it against 7 state-of-the-art MoE serving systems, achieving up to 6.68$\times$ MoE layer speedup and a geometric mean speedup of 1.86$\times$ over the fastest competing baseline.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Distributed Serving
Load Balancing
Expert Parallelism
Bottleneck Communication
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Distributed Serving
Load Balancing
Communication Optimization
Expert Parallelism
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
C
Chang Chen
UCSD
A
Andrew Yang
Stanford University
T
Tiancheng Chen
ETH ZΓΌrich
Jiangfei Duan
Jiangfei Duan
The Chinese University of Hong Kong
Machine Learning Systems
X
Xinwei Qiang
UCSD
Z
Zhongkai Yu
UCSD
X
Xiang Fang
UCSD
Y
Yufei Ding
UCSD