moe routing

Design, implement, and analyze routing mechanisms and control strategies for mixture-of-experts (MoE) models that decide which expert(s) receive each input, including steering signals, router behavior diagnostics, and domain-aware routing policies. This includes creating single- and multi-expert selection rules and coarse-to-fine classifiers, losses and regularizers to shape router probabilities and margins, aggregation methods for expert outputs, and analyses that trade off accuracy, throughput, and training/gradient/communication effects.

moerouting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study investigates the relationship between safety behaviors and expert routing mechanisms in aligned Mixture-of-Experts (MoE) large language models. The authors find that a model’s safety capabilities can be concentrated in a small subset of experts and are largely independent of the routing policy, rather than being driven by dedicated refusal-oriented experts. To leverage this insight, they propose the Router-Agnostic Safety-critical Expert Tuning (RASET) framework, which integrates a contrastive routing sensitivity criterion with parameter-efficient fine-tuning to precisely identify and optimize safety-critical experts without altering the original routing behavior. Experiments demonstrate that RASET significantly steers model safety outputs with minimal semantic interference, revealing for the first time the existence of localized, manipulable expert-level safety mechanisms—and their potential vulnerabilities—within MoE architectures.

expert specializationMixture-of-Expertsrouter-driven activation

This work addresses a key limitation in sparse Mixture-of-Experts (MoE) models, where the routing mechanism jointly handles expert selection and output weighting, potentially constraining performance. The study provides the first systematic validation that these two functions should be decoupled and introduces Fixed Dispatch with Adaptive Aggregation (FDAA): a lightweight, learnable aggregation head is added atop a frozen backbone and fixed expert assignments, enabling end-to-end optimization of aggregation weights via the language modeling objective. Evaluated on pretrained MoE models such as OLMoE and DeepSeek-V2-Lite, FDAA achieves a 0.1523 reduction in cross-entropy on WikiText-103 and demonstrates consistent improvements across diverse benchmarks including C4 and PTB, confirming both the efficacy and generality of the proposed decoupling strategy.

expert aggregationexpert dispatchMixture-of-Experts

This work addresses the lack of effective design principles for routers in existing Mixture-of-Experts (MoE) models, which struggle to accurately capture the affinity between tokens and experts. The authors propose, for the first time, using the dominant singular directions of expert matrices as the target for router design and introduce a novel “power iteration followed by shrinkage” paradigm. During pretraining, they employ manifold optimization to dynamically align the router’s row vectors with these dominant singular directions. This approach achieves a favorable balance among alignment accuracy, computational efficiency, and training stability. Experiments on MoE models ranging from 1B to 11B parameters demonstrate substantial performance improvements, validating the effectiveness of the proposed router redesign strategy.

expert representationMixture-of-Expertsrouter design

This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.

dynamic routingexpert capacityMixture-of-Experts

Latest Papers

What's happening recently
View more

This study investigates the dynamic origins of expert load imbalance in Mixture-of-Experts (MoE) routing. By constructing a mean-field limit dynamical model for two-expert Softmax routing, the authors uncover the adaptive mechanisms and load evolution patterns inherent in the system. Theoretical analysis reveals that under symmetric conditions, the system undergoes a supercritical pitchfork bifurcation, while the introduction of external asymmetry induces a cusp catastrophe structure, offering a low-dimensional, controllable explanation for abrupt load imbalances. Leveraging bifurcation theory and the canonical form of the cusp catastrophe, the authors derive an exact parametric equation for the bifurcation set. Experimental validation using PyTorch with hard top-1 routing successfully reproduces the abrupt load shifts observed in real MoE systems, confirming the theoretical predictions.

BifurcationCusp CatastropheLoad Imbalance

This work addresses the suboptimality of standard top-k routing in Mixture-of-Experts language models for complex reasoning tasks, where direct evaluation of routing efficacy has been lacking. By freezing model parameters and comparing standard routing against counterfactual alternatives with equivalent computational cost—using next-token prediction probabilities along ground-truth reasoning trajectories as a utility metric—the study reveals that routers perform well on high-confidence tokens but fail at fragile reasoning steps. This limitation stems from training objectives that optimize only the executed path and rely on statistical load balancing. To mitigate this, the authors propose fine-tuning only the final-layer router, which significantly improves pass@K performance on AIME 2024+2025 and HMMT 2025 benchmarks in Qwen3-30B-A3B and GPT-OSS-20B.

expert allocationlanguage modelsMixture-of-Experts

This study investigates how the routing mechanism of the Mixtral 8x7B-Instruct model influences safety outcomes in response to both benign and harmful prompts. By jointly analyzing expert activation frequencies and router gating gradients—and integrating targeted expert suppression with cross-group expert categorization—the work reveals, for the first time, the deep dependency and distributed nature of safety-related routing decisions. The findings demonstrate that safety-critical experts are broadly dispersed yet concentrated in specific layers; moreover, selectively suppressing experts identified via gradient-based importance significantly reduces restricted responses while inducing fewer side effects, thereby overcoming the limitations of single-metric analyses.

harmful promptslanguage modelsmixture-of-experts

This work addresses the inefficiency of conventional Mixture-of-Experts (MoE) models, which employ a fixed top-k expert selection strategy that fails to dynamically allocate computational resources according to individual token demands. To overcome this limitation, the authors propose a training-free, plug-in method for inference that introduces, for the first time in MoE architectures, an elbow-point detection mechanism. By analyzing the probability distribution output by the router, this approach adaptively determines the number of experts to activate per token. Integrating principles from ranking and load balancing theory, the method achieves dynamic resource allocation while preserving balanced expert utilization. Experimental results demonstrate that the proposed technique reduces average inference latency by 5.3% across mainstream MoE models without compromising accuracy on six benchmark evaluations.

compute allocationdynamic routingexpert selection

This study addresses the limitations of flat routing in conventional Mixture-of-Experts (MoE) models, which often suffer from expert load imbalance and a lack of topological structure. To overcome these issues, this work proposes a hierarchical binary decision tree routing architecture that dynamically estimates branch probabilities via exponential moving averages to balance traffic across subtrees. This mechanism achieves load balancing without auxiliary losses while theoretically preventing routing collapse. Experimental results demonstrate that the proposed method preserves task accuracy across multiple benchmark datasets while significantly reducing cross-device communication overhead in distributed settings, thereby enabling efficient and balanced expert utilization.

Communication OverheadExpert Utilization ImbalanceFlat Routing

Hot Scholars

RW

Rongzhe Wei

Georgia Institute of Technology
LLM TrustworthyGraph AnalysisPrivacyUnlearning
IB

Irina Belousova

AI/ML Applied Research Scientist & Manager
Deep LearningMachine LearningNatural Language Processing
YC

Yixin Cao

Fudan University
Natural Language ProcessingKnowledge EngineeringMulti-modal data processing
XZ

Xiaowei Zhu

Ant Research
Graph DatabaseBig Data SystemsPrivacy-Preserving ComputationAI Infra