CIPHER-MoE: Balancing Efficiency and Routing Fidelity in Trillion-Scale MoE Training

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the severe expert load imbalance caused by uneven token routing in trillion-parameter Mixture-of-Experts (MoE) training, which significantly limits training efficiency and hardware utilization. To mitigate this issue, we propose an affinity-aware expert filtering mechanism coupled with explicit capacity control, effectively alleviating load imbalance while preserving the original router selection policy. Notably, this approach incurs minimal system overhead, requiring neither additional hardware nor complex runtime designs. Experimental results demonstrate that the proposed method reduces top-1 expert load by 64.9% and achieves a 1.1× to 1.94× training speedup without compromising model quality.
📝 Abstract
Mixture-of-Experts (MoE) has been widely adopted in recent large language model (LLM) architectures. However, scaling up MoE in LLM training introduces system-level challenges on training, where non-uniform token routing can lead to highly imbalanced workloads across experts and devices, further destabilizing the training process. With trillion-scale LLMs, imbalanced expert workloads further amplify the resource cost of MoE training, resulting in degraded training efficiency and hardware utilization for underloaded experts, while hot experts require additional resources to accommodate excessive workloads. Recent studies address imbalanced MoE training through intricate parallelism strategies or resource reallocation. However, these system-level approaches often introduce additional resource requirements and considerable orchestration complexity, which become increasingly difficult to afford when training trillion-parameter LLMs under constrained computational resources. This work introduces CIPHER-MoE, which mitigates MoE workload imbalance while keeping the router's token-side Top-K selection unchanged. CIPHER-MoE applies affinity-aware Expert-to-Token filtering with explicit capacity control to reduce hotspot expert workloads without additional hardware resources or complex runtime design. The proposed method has been evaluated on large-scale MoE models, including DeepSeek-V4-Pro, showing up to 64.9 percentage points Top-1 expert workload reduction and 1.10$\times$-1.94$\times$ training acceleration, while preserving the training quality. The source code will be released soon.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
workload imbalance
trillion-scale LLM training
training efficiency
hardware utilization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts (MoE)
Workload Balancing
Expert-to-Token Filtering
Trillion-Scale LLM Training
Capacity Control
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.