Function Over Form: Distributional Orthogonalization in Mixture-of-Experts with Replica Expert Mechanism

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the underutilization of model capacity in Mixture-of-Experts (MoE) architectures caused by expert collapse and representation redundancy. To mitigate these issues, it proposes a distribution orthogonalization loss alongside a replica expert mechanism, shifting the optimization perspective from static weight diversity to dynamic routing behavior constraints that promote functional specialization by penalizing routing overlap. Furthermore, this work designs a bi-level load balancing algorithm integrating global batch adjustment with real-time token dispatch, augmented by high-dimensional binary load signatures for optimized scheduling. Extensive evaluations on 4.8B and 30B MoE models demonstrate that the proposed approach significantly outperforms existing routing algorithms across downstream tasks while maintaining comparable training efficiency, thereby effectively enhancing overall system performance.
📝 Abstract
The scaling of LLMs increasingly relies on MoE architectures to decouple active computation from total parameter count. However, the efficacy of MoE is often constrained by expert collapse and representation redundancy, both leading to underutilization of model capacity. To address these challenges, this paper proposes Distributional Orthogonalization Loss (DO-loss), an auxiliary regularization that shifts the focus from static weight diversity to dynamic routing behavior. By representing each expert's token assignment history as a high-dimensional binary load signature, DO-loss penalizes signature overlap to prevent expert collapse while encouraging functional specialization. To align this algorithmic design with system efficiency, we further introduce the Replica Expert Mechanism (REM), which improves load balancing through a two-tiered strategy: adjusting replica expert placement at the global-batch level and performing real-time token dispatching at the micro-batch level. Empirical evaluations demonstrate that our method outperforms the evaluated routing algorithms on downstream tasks for both 4.8BA0.5B and 30BA3B MoE models, while maintaining comparable training efficiency.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Expert Collapse
Representation Redundancy
Load Balancing
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distributional Orthogonalization Loss
Replica Expert Mechanism
Mixture-of-Experts
Expert Collapse
Load Balancing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jinfan He
School of EECS, Peking University, China
Y
Yunzhuo Liu
Tencent Hunyuan, China
K
Kai Zhang
Tencent Hunyuan, China
Weidong Han
Weidong Han
Tencent Inc., School of Data Science, Fudan University
Large Language ModelNLPMulti-Modal
K
Key
Tencent Hunyuan, China
R
Rayying
Tencent Hunyuan, China