🤖 AI Summary
This study addresses the substantial All-to-All communication overhead in Mixture-of-Experts (MoE) expert parallelism, which accounts for 45%–60% of training step time. By observing significant intra-layer and inter-layer correlations in expert selection during early pre-training stages, this work proposes a routing-correlation-based communication optimization method. Specifically, correlation-aware expert placement strategies and token shuffling mechanisms are designed within the Megatron-LM framework to effectively reduce cross-GPU data transfers. Experimental results demonstrate that the proposed approach reduces All-to-All communication latency by 1.16× to 2.63× and achieves up to a 1.41× end-to-end training step speedup, substantially improving the distributed training efficiency of MoE models.
📝 Abstract
Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts. Under expert parallelism (EP) the experts are distributed across GPUs, and every MoE layer runs all-to-all collectives in the forward and backward passes to dispatch tokens to their experts and then combine the results. On a cluster with 8 AMD Instinct MI300X GPUs per node, these collectives can take 45% of the training step at EP32 with top-2 routing and 60% with top-6 routing. We find that early in pretraining routers have already learned to assign tokens to experts in correlated patterns, both within a layer and across layers. At top-2, 0.8% of the expert pairs in a layer are selected together by 42% of tokens, and the experts a token selects at one layer predict the experts it selects at the next layer. We use these correlations to keep more token--expert assignments on the token's own GPU, which reduces communication across GPUs and across nodes. Correlated expert placement puts experts that are often selected together on the same GPU. Combined with a dispatcher that sends each token to each GPU once, it removes up to 58% of dispatched rows. Token shuffling applies when sequence parallelism shards tokens across the EP group. It moves each token to the GPU predicted to hold its next-layer experts during the reduce-scatter that follows attention. On one node this raises the share of token--expert assignments served on the token's GPU from 12.5% to 59%. In Megatron-LM, across EP degrees from 8 to 64 with top-2 and top-6 routing, the two methods reduce all-to-all time by 1.16-2.63X and end-to-end step time by up to 1.41X. Neither method changes the models' underlying routing decisions or expert parameters.