ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the training instability and inefficiency caused by expert load imbalance in extremely sparse Mixture-of-Experts (MoE) models. To this end, it proposes ID Balancing, a method that unifies existing auxiliary losses within a Proportional-Integral-Derivative (PID) control framework. By introducing an integral term scaled by error magnitude and a derivative term activated only upon performance degradation, the approach dynamically adjusts routing weights to achieve auxiliary-loss-free load balancing. Experimental results demonstrate that under a Top-3 routing configuration, the maximum violation rate is reduced by over 50%. Furthermore, when scaling model parameters to 69.9B, the proposed method maintains stable training while achieving performance superior to 89.6% of the baseline configurations, highlighting its effectiveness and scalability for large-scale MoE architectures.
📝 Abstract
Scaling Large Language Models (LLMs) via Mixture-of-Experts (MoE) enables massive parameter growth with nearly constant per-token computation. However, further scaling the parameter count requires increasingly sparse routing, where expert load imbalance becomes more severe. This imbalance reduces parameter utilization and training efficiency, and can undermine training stability, becoming a bottleneck to reliable scaling. In this work, we unify two representative auxiliary-loss-free methods as incomplete Proportional-Integral-Derivative (PID) controllers: DeepSeek's loss-free method acts as a fixed-step integral controller, while Kimi K3's Quantile Balancing functions as a generalized proportional controller. Building on this control perspective, we propose ID Balancing, an Integral-Derivative controller. It scales its integral term with load error and activates its derivative term only when imbalance worsens, enabling stronger corrections for large or worsening errors and smaller updates near balance. Evaluated across Top-$10$, Top-$5$, and Top-$3$ routing over $768$ experts, ID Balancing reduces worst-case backbone MaxVio and training-average backbone MinVio by over $50\%$ and $12\%$, respectively, relative to the best baselines in the Top-$3$ setting. When the total parameter count increases from $18.9$B to $69.9$B (Top-$10$-of-$768$), ID Balancing's worst-case backbone MaxVio remains nearly unchanged and is approximately $89.6\%$ lower than that of the auxiliary-loss baseline. ID Balancing also maintains competitive language-modeling and downstream performance. The advantages of ID Balancing grow as sparsity increases, making it a promising solution for scaling larger, sparser MoE models.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Load Imbalance
Sparse Routing
Training Stability
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
PID Controller
Load Balancing
Sparse Routing
Training Stability
🔎 Similar Papers
No similar papers found.
P
Peng Jin
Qwen Team, Alibaba Token Hub, Alibaba Group
Zihan Qiu
Zihan Qiu
Qwen Team, Alibaba Group & IIIS, Tsinghua University
Mixture of ExpertsModular Deep LearningInterpretability
Z
Zekun Wang
Qwen Team, Alibaba Token Hub, Alibaba Group
B
Bo Zheng
Qwen Team, Alibaba Token Hub, Alibaba Group
Y
Yang Xu
Qwen Team, Alibaba Token Hub, Alibaba Group
T
Tian Xie
Qwen Team, Alibaba Token Hub, Alibaba Group
X
Xiao Li
Qwen Team, Alibaba Token Hub, Alibaba Group
H
Huaqing Zhang
Qwen Team, Alibaba Token Hub, Alibaba Group
Haoran Lian
Haoran Lian
Beihang University
Natural Language Processing
Rui Men
Rui Men
Qwen Team, Alibaba Group & Peking University
NLP
D
Dayiheng Liu
Qwen Team, Alibaba Token Hub, Alibaba Group