MegaFlux: Skew-Resilient MoE Megakernels via Pipelined Expert Replication

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the routing skew caused by fixed expert placement in Mixture-of-Experts (MoE) models, which leads to imbalanced GPU workloads and communication latency bottlenecks. To mitigate this, we propose a runtime dynamic replication strategy for hotspot experts and integrate pipelined communication into persistent execution kernels, enabling deep overlap of forward and backward computation with communication to eliminate idle waiting. Building upon CuTeDSL, we extend TensorRT-LLM by introducing the first runtime expert replication decision mechanism and developing novel MoE forward and backward megakernels. Evaluated on a B200 cluster, our approach achieves speedups of 1.45× and 1.28× for the forward and backward passes, respectively, and delivers up to a 1.26× improvement in end-to-end inference throughput.
📝 Abstract
Mixture-of-experts (MoE) megakernels fuse expert-parallel communication with expert computation. However, under fixed expert placement, routing skew creates GPU stragglers: overloaded GPUs determine layer latency while others sit idle. Replicating hot experts can shift work to underloaded GPUs, but dynamic replicas introduce additional work: replicas must receive expert weights to execute and, during training, their partial weight gradients must be reduced at the expert owners. We present MegaFlux, which makes expert replication a runtime decision and pipelines the communication induced by replication within persistent MoE execution. An on-device planner jointly selects replica locations and assigns tile-aligned token blocks under a per-GPU replica budget, leaving router outputs unchanged. The forward and backward megakernels realize pipelined expert replication: replicas begin computation as their required weights arrive, while backward overlaps replica-gradient reduction with ongoing expert computation. MegaFlux extends TensorRT-LLM's CuTeDSL MegaMoE forward kernel and introduces a new backward MoE megakernel. Across 147 configurations per direction on eight NVIDIA B200 GPUs, MegaFlux achieves geometric-mean speedups of $1.45\times$ for forward and $1.28\times$ for backward over the same megakernels with fixed placement, peaking at $2.14\times$ and $2.64\times$. In ablations, pipelining hides $56$--$76$% of replica-weight transfer cost in forward and $91$--$100$% of combined weight-transfer and replica-gradient-reduction cost in backward, yielding up to $13.2$% and $26.7$% additional layer-latency reductions over the same replication plans with these operations executed separately. Integrated into vLLM for DeepSeek-V4-Pro prefill, MegaFlux delivers $1.13$--$1.26\times$ median end-to-end speedups over fixed placement.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
routing skew
GPU stragglers
expert replication
load imbalance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Megakernel
Expert Replication
Pipelined Communication
Skew-Resilient
🔎 Similar Papers
No similar papers found.