Score
Design, implement, and analyze routing mechanisms and control strategies for mixture-of-experts (MoE) models that decide which expert(s) receive each input, including steering signals, router behavior diagnostics, and domain-aware routing policies. This includes creating single- and multi-expert selection rules and coarse-to-fine classifiers, losses and regularizers to shape router probabilities and margins, aggregation methods for expert outputs, and analyses that trade off accuracy, throughput, and training/gradient/communication effects.
To address cross-layer optimization challenges—including low inference speed, high energy consumption, and poor hardware utilization—in large-scale Mixture-of-Experts (MoE) model deployment, this paper proposes the first unified taxonomy spanning model, system, and hardware stacks. Methodologically, it integrates expert pruning, dynamic routing with load balancing, multi-granularity compression (pruning, quantization, distillation), distributed scheduling, and hardware-aware compilation. We innovatively establish a cross-layer co-optimization framework and open-source an actively maintained *Awesome MoE Inference* knowledge repository. Our contributions include a structured technical landscape that clarifies core challenges and evolutionary trends, significantly improving inference efficiency and energy efficiency—enabling low-latency, high-throughput, and power-efficient industrial MoE deployment.
This study investigates the relationship between safety behaviors and expert routing mechanisms in aligned Mixture-of-Experts (MoE) large language models. The authors find that a model’s safety capabilities can be concentrated in a small subset of experts and are largely independent of the routing policy, rather than being driven by dedicated refusal-oriented experts. To leverage this insight, they propose the Router-Agnostic Safety-critical Expert Tuning (RASET) framework, which integrates a contrastive routing sensitivity criterion with parameter-efficient fine-tuning to precisely identify and optimize safety-critical experts without altering the original routing behavior. Experiments demonstrate that RASET significantly steers model safety outputs with minimal semantic interference, revealing for the first time the existence of localized, manipulable expert-level safety mechanisms—and their potential vulnerabilities—within MoE architectures.
This work addresses a key limitation in sparse Mixture-of-Experts (MoE) models, where the routing mechanism jointly handles expert selection and output weighting, potentially constraining performance. The study provides the first systematic validation that these two functions should be decoupled and introduces Fixed Dispatch with Adaptive Aggregation (FDAA): a lightweight, learnable aggregation head is added atop a frozen backbone and fixed expert assignments, enabling end-to-end optimization of aggregation weights via the language modeling objective. Evaluated on pretrained MoE models such as OLMoE and DeepSeek-V2-Lite, FDAA achieves a 0.1523 reduction in cross-entropy on WikiText-103 and demonstrates consistent improvements across diverse benchmarks including C4 and PTB, confirming both the efficacy and generality of the proposed decoupling strategy.
This work addresses the lack of effective design principles for routers in existing Mixture-of-Experts (MoE) models, which struggle to accurately capture the affinity between tokens and experts. The authors propose, for the first time, using the dominant singular directions of expert matrices as the target for router design and introduce a novel “power iteration followed by shrinkage” paradigm. During pretraining, they employ manifold optimization to dynamically align the router’s row vectors with these dominant singular directions. This approach achieves a favorable balance among alignment accuracy, computational efficiency, and training stability. Experiments on MoE models ranging from 1B to 11B parameters demonstrate substantial performance improvements, validating the effectiveness of the proposed router redesign strategy.
本文通过局部聚合视角分析混合专家模型,探讨路由、稀疏激活和共享专家等设计选择的统计作用,分离出逼近误差、专家学习误差和路由器估计误差。
This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.
This study investigates the dynamic origins of expert load imbalance in Mixture-of-Experts (MoE) routing. By constructing a mean-field limit dynamical model for two-expert Softmax routing, the authors uncover the adaptive mechanisms and load evolution patterns inherent in the system. Theoretical analysis reveals that under symmetric conditions, the system undergoes a supercritical pitchfork bifurcation, while the introduction of external asymmetry induces a cusp catastrophe structure, offering a low-dimensional, controllable explanation for abrupt load imbalances. Leveraging bifurcation theory and the canonical form of the cusp catastrophe, the authors derive an exact parametric equation for the bifurcation set. Experimental validation using PyTorch with hard top-1 routing successfully reproduces the abrupt load shifts observed in real MoE systems, confirming the theoretical predictions.
This work addresses the suboptimality of standard top-k routing in Mixture-of-Experts language models for complex reasoning tasks, where direct evaluation of routing efficacy has been lacking. By freezing model parameters and comparing standard routing against counterfactual alternatives with equivalent computational cost—using next-token prediction probabilities along ground-truth reasoning trajectories as a utility metric—the study reveals that routers perform well on high-confidence tokens but fail at fragile reasoning steps. This limitation stems from training objectives that optimize only the executed path and rely on statistical load balancing. To mitigate this, the authors propose fine-tuning only the final-layer router, which significantly improves pass@K performance on AIME 2024+2025 and HMMT 2025 benchmarks in Qwen3-30B-A3B and GPT-OSS-20B.
This study investigates how the routing mechanism of the Mixtral 8x7B-Instruct model influences safety outcomes in response to both benign and harmful prompts. By jointly analyzing expert activation frequencies and router gating gradients—and integrating targeted expert suppression with cross-group expert categorization—the work reveals, for the first time, the deep dependency and distributed nature of safety-related routing decisions. The findings demonstrate that safety-critical experts are broadly dispersed yet concentrated in specific layers; moreover, selectively suppressing experts identified via gradient-based importance significantly reduces restricted responses while inducing fewer side effects, thereby overcoming the limitations of single-metric analyses.
This work addresses the inefficiency of conventional Mixture-of-Experts (MoE) models, which employ a fixed top-k expert selection strategy that fails to dynamically allocate computational resources according to individual token demands. To overcome this limitation, the authors propose a training-free, plug-in method for inference that introduces, for the first time in MoE architectures, an elbow-point detection mechanism. By analyzing the probability distribution output by the router, this approach adaptively determines the number of experts to activate per token. Integrating principles from ranking and load balancing theory, the method achieves dynamic resource allocation while preserving balanced expert utilization. Experimental results demonstrate that the proposed technique reduces average inference latency by 5.3% across mainstream MoE models without compromising accuracy on six benchmark evaluations.
This study addresses the limitations of flat routing in conventional Mixture-of-Experts (MoE) models, which often suffer from expert load imbalance and a lack of topological structure. To overcome these issues, this work proposes a hierarchical binary decision tree routing architecture that dynamically estimates branch probabilities via exponential moving averages to balance traffic across subtrees. This mechanism achieves load balancing without auxiliary losses while theoretically preventing routing collapse. Experimental results demonstrate that the proposed method preserves task accuracy across multiple benchmark datasets while significantly reducing cross-device communication overhead in distributed settings, thereby enabling efficient and balanced expert utilization.