Cobalt: Leveraging Expert Co-activation for Efficient Distributed MoE Training

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the efficiency bottlenecks in distributed Mixture-of-Experts (MoE) training caused by limited inter-node bandwidth and imbalanced expert workloads. To overcome these challenges, this work proposes an adaptive placement strategy leveraging expert co-activation characteristics, which synergistically optimizes expert parallelism through two-stage placement planning and communication-aware routing. The primary contribution lies in being the first to integrate co-activation patterns into dynamic placement and task allocation, thereby significantly reducing inter-node traffic while balancing computational loads. Experimental evaluations on 32 NVIDIA B200 GPUs demonstrate that the proposed approach achieves up to a 2.41× speedup and reduces cross-node communication volume by over 75%, highlighting its effectiveness for large-scale MoE training.
📝 Abstract
Mixture-of-Experts (MoE) has increasingly become a mainstream approach for scaling large language models, as it expands model capacity while keeping computation cost nearly constant. Training large-scale MoE models relies on Expert Parallelism (EP), which distributes expert replicas across GPUs and exchanges tokens through all-to-all communication. The efficiency of EP is often constrained by two system bottlenecks: cross-node token transfers are limited by inter-node bandwidth, while skewed expert workloads lead to imbalanced computation across GPUs. Prior work mitigates these bottlenecks based on per-expert workload statistics, but overlooks the fact that experts could share the communication. In this work, we empirically present the observation that many pairs of experts are frequently co-activated by individual tokens. Motivated by this, we present Cobalt, an efficient MoE training framework that leverages expert co-activation to reduce cross-node traffic and workload imbalance. Cobalt adopts a two-stage expert layout planner that adapts expert layout to the evolving expert co-activation and workload conditions. It periodically co-locates frequently co-activated experts on the same node to reduce the cross-node communication, and performs per-step intra-node adjustment to rebalance the workloads. Subsequently, we develop a communication-aware task assignment method that routes tokens to fewer remote nodes based on the current expert layout. Experiments on 32 B200 GPUs show that Cobalt achieves up to 1.53-2.41 times (1.28-1.89 times on average) of speedup compared to existing MoE training frameworks, while reducing cross-node token traffic by 75.74%-99.26%.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Expert Parallelism
distributed training
communication bottleneck
workload imbalance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Expert Co-activation
Distributed Training
Expert Layout Planner
Communication-aware Task Assignment