🤖 AI Summary
This study addresses the underutilization of model capacity in Mixture-of-Experts (MoE) architectures caused by expert collapse and representation redundancy. To mitigate these issues, it proposes a distribution orthogonalization loss alongside a replica expert mechanism, shifting the optimization perspective from static weight diversity to dynamic routing behavior constraints that promote functional specialization by penalizing routing overlap. Furthermore, this work designs a bi-level load balancing algorithm integrating global batch adjustment with real-time token dispatch, augmented by high-dimensional binary load signatures for optimized scheduling. Extensive evaluations on 4.8B and 30B MoE models demonstrate that the proposed approach significantly outperforms existing routing algorithms across downstream tasks while maintaining comparable training efficiency, thereby effectively enhancing overall system performance.
📝 Abstract
The scaling of LLMs increasingly relies on MoE architectures to decouple active computation from total parameter count. However, the efficacy of MoE is often constrained by expert collapse and representation redundancy, both leading to underutilization of model capacity. To address these challenges, this paper proposes Distributional Orthogonalization Loss (DO-loss), an auxiliary regularization that shifts the focus from static weight diversity to dynamic routing behavior. By representing each expert's token assignment history as a high-dimensional binary load signature, DO-loss penalizes signature overlap to prevent expert collapse while encouraging functional specialization. To align this algorithmic design with system efficiency, we further introduce the Replica Expert Mechanism (REM), which improves load balancing through a two-tiered strategy: adjusting replica expert placement at the global-batch level and performing real-time token dispatching at the micro-batch level. Empirical evaluations demonstrate that our method outperforms the evaluated routing algorithms on downstream tasks for both 4.8BA0.5B and 30BA3B MoE models, while maintaining comparable training efficiency.