π€ AI Summary
This study addresses the GPU memory bottleneck caused by the all-to-all exchange mechanism of Top-k routing in distributed Mixture-of-Experts (MoE) training. We propose RelayMoE, which introduces a novel ring-streaming execution architecture that avoids constructing full buffers by performing local computation over experts or tokens. By integrating dynamic routing strategies, asynchronous communication-computation overlap, and memory-aware recomputation techniques, RelayMoE achieves highly memory-efficient optimization. Experimental results demonstrate that this method accelerates single-layer computation by 2Γ and improves overall throughput by 2.02Γ, while enabling a 2.85Γ extension in sequence length. These findings indicate that RelayMoE effectively overcomes the memory wall limitations inherent in large-scale MoE training.
π Abstract
As Mixture-of-Experts (MoE) models scale toward hundreds of experts and higher top-$k$ routing, memory efficiency in distributed training becomes a critical bottleneck. Peak memory is dominated by the MoE block, not attention: every intermediate buffer in the MoE dispatch pipeline is individually scaled by top-k routing. The standard all-to-all dispatcher sends all routed tokens in a single collective step, requiring the full top-$k$-expanded buffer to be constructed at once. In this work, we propose RelayMoE, a ring-based MoE execution model that computes locally as expert weights or tokens circulate, avoiding full top-$k$-expanded dispatch buffers. RelayMoE selects between expert and token routing according to communication volume and overlaps transfers with computation. The ring structure naturally supports memory-efficient MoE recomputation during backward: each hop reconstructs expert intermediates, uses them to compute gradients, and releases them before the next hop. The saved memory supports longer sequences and larger batches, or retains more attention activations to reduce attention recomputation and improve training throughput. We evaluate RelayMoE on 30B$-$57B production MoE models and varied expert configurations. In single-layer MoE experiments, RelayMoE achieves a $2\times$ average speedup over Megatron-LM. In full-model training under the same GPU memory budget, it improves throughput by up to $2.02\times$ and extends the largest tested trainable sequence length by up to $2.85\times$.