🤖 AI Summary
This study addresses the degradation of end-to-end performance in Mixture-of-Experts (MoE) serving caused by imbalanced expert loads. To this end, it proposes AFORE, a system built upon the Attention-FFN Decoupling (AFD) architecture. AFORE introduces a novel near-future demand-aware dynamic expert migration mechanism, integrated with micro-batch scheduling and pipeline overlapping, to fully conceal migration latency within the AFD pipeline window. Additionally, the system leverages high-speed NVLink inter-GPU communication alongside lightweight demand prefetching. Experimental results demonstrate that, compared to static placement baselines, AFORE improves throughput by 10.1%–17.6% and reduces P95 latency by 7.1%–9.5%.
📝 Abstract
Efficient serving of Mixture-of-Experts (MoE) models is challenging due to large expert parameters, input-dependent expert activation, and dynamic workloads. Expert parallelism distributes expert computation across GPUs, while attention-FFN disaggregation (AFD) separates attention and feed-forward computation into independent worker pools. However, we observe that a naive AFD implementation could make expert load imbalance more harmful: once FFN computation becomes an independent pipeline stage, overloaded experts directly slow the FFN stage and degrade end-to-end serving performance. To solve this problem, we present AFORE, a timely expert reconfiguration system for AFD-based MoE serving. AFORE exploits two architectural properties of AFD. First, AFD exposes the expert-token distribution of upcoming microbatches before they reach FFN execution, enabling placement decisions based on near-future demand instead of stale historical profiles. Second, AFD creates a pipeline window in which expert migration for a target microbatch can be overlapped with the computation of preceding in-flight microbatches. AFORE formulates expert reconfiguration as a microbatch-aware scheduling problem and uses a migration-aware scheduler to decide when and which experts to migrate. AFORE further implements lightweight demand prefetching and NVLink-based GPU-GPU expert migration to reduce reconfiguration overhead. Evaluation on a 110B-parameter MoE model across four dynamic workloads shows that AFORE improves output throughput by 10.1-17.6% and reduces P95 inter-token latency by 7.1-9.5% compared with the strongest competing baseline. Compared with static placement, AFORE improves throughput by 29.8% on average and reduces P95 inter-token latency by 18.2% on average. Migration profiling further shows that AFD pipeline overlap can fully hide expert-migration latency.