TopoEP: Topology-Aware Load Balancing for Expert-Parallel MoE Training

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the load imbalance and CPU scheduling bottlenecks induced by dynamic routing during Mixture-of-Experts (MoE) training. To mitigate these issues, this work proposes a GPU-native, topology-aware load balancing system. The method introduces a deterministic GPU-based solver that enables parallel planning of inter-node placement and intra-node refinement, thereby completely eliminating host-side synchronization overhead. The proposed system is integrated into the Megatron-LM framework and validated on an NVIDIA H800 cluster. Experimental results demonstrate that, on a 32-GPU cluster, the end-to-end training throughput improves by 6.2% to 11.4%. Consequently, this approach establishes a new paradigm for the efficient large-scale training of MoE models.
📝 Abstract
Dynamic routing creates severe load imbalance in large-scale expert-parallel Mixture-of-Experts (MoE) training, turning GPUs that host hot experts into stragglers. As each MoE layer waits for its slowest rank, these stragglers prolong the expert-parallel stage and reduce overall training efficiency. Existing expert-parallelism load-balancing (EPLB) systems commonly compute load-balancing plans on the CPU, incurring device--host data transfers and cross-rank synchronization that make scheduling at every layer and microbatch expensive. Their planning formulations also overlook the hierarchical communication costs of modern scale-up and scale-out GPU clusters. We present \textit{TopoEP}, a GPU-native, topology-aware load-balancing system for large-scale MoE training. At each MoE layer and training microbatch, \textit{TopoEP} converts the current routing result into hot-expert replication and token-rerouting decisions and executes the resulting plan without data-dependent host synchronization, reducing critical-path overhead. To generate these decisions, \textit{TopoEP} uses a deterministic GPU solver that performs inter-node placement followed by intra-node refinement, allowing all ranks to independently produce bitwise-identical plans. On a 32-GPU NVIDIA H800 cluster, integrating \textit{TopoEP} with Megatron-LM improves end-to-end training throughput by 6.2\%--11.4\% across three representative MoE models.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
load imbalance
expert parallelism
dynamic routing
hierarchical communication cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Load Balancing
Topology-Aware
GPU-Native
Expert Parallelism
🔎 Similar Papers
No similar papers found.