Three Phases of Expert Routing: How Load Balance Evolves During Mixture-of-Experts Training

📅 2026-04-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work uncovers the dynamic trade-off between load balancing and model quality in Mixture-of-Experts (MoE) training. By modeling token routing as a congestion game and introducing an effective congestion coefficient γ_eff to monitor the entire training trajectory, the study reveals—for the first time—a non-monotonic three-phase evolution of load balancing: surge, stabilization, and relaxation. This pattern elucidates the intrinsic mechanism of prioritizing balancing early in training and model quality later. The approach integrates temperature-scaled softmax, multi-type congestion decomposition, token clustering, and range diagnostics (K/M, ε_l). Evaluated on OLMoE-1B–7B and OpenMoE-8B, it achieves a 30% average improvement in load prediction, model quality estimation consistency with Pearson correlation r ≥ 0.89, and L1 error approaching the theoretical lower bound.

Technology Category

Machine Learning: Mixture of Experts (MoE)Search and Optimization: Learning to SearchPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

User Modeling, Personalization and Recommendation: Practical large-scale studies of user experienceGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
We model Mixture-of-Experts (MoE) token routing as a congestion game with a single effective parameter, the congestion coefficient gamma_eff, that quantifies the balance-quality tradeoff. Tracking gamma_eff across training checkpoints of two open-source MoE models, OLMoE-1B-7B (20 checkpoints, with dense sampling in the surge region) and OpenMoE-8B (6 checkpoints), reveals a three-phase trajectory: a surge phase where the router learns to balance load (gamma_eff: 14 to 36-39, peaking in the step 30K-40K region), a stabilization phase where experts specialize under steady balance (B_0: 2.4 to 2.3, steps 100K-400K), and a relaxation phase where the router trades balance for quality as experts differentiate (gamma_eff: 27 to 9, steps 400K-1.2M). This non-monotone trajectory, invisible to post-hoc analysis of converged models, reveals that early MoE training prioritizes balance while late training prioritizes quality. The theoretical framework is honest about its limits: the single-type equilibrium reduces to temperature-scaled softmax (held-out L1: MFG = 0.199 vs. softmax = 0.200). The game is not a better predictor; it reveals what the temperature means and, critically, how that temperature evolves. We complement the dynamics with an effective congestion decomposition, a multi-type extension that improves load prediction via token clustering on all 16 layers (mean: 30%), scope diagnostics (K/M, epsilon_l), and robustness verification across four independent quality estimators (r >= 0.89). All confidence intervals are from bootstrap resampling over 50 independent text batches.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
expert routing
load balancing
training dynamics
congestion game
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
congestion game
load balancing
training dynamics
routing optimization
C
Charafeddine Mouzouni
OPIT – Open Institute of Technology, and Cohorte AI, Paris, France