mixture-of-experts routing

Designs and implements mechanisms that route inputs, internal activations, or contextual signals to a subset of specialist model components (‘experts’) — covering conditional, class- or task-conditioned, dynamic, state-gated, frequency- or residual-guided, scale-adaptive, sparse vs dense, bi-modal or autoencoder-based routing and bridge/gating variants — so different experts handle different subproblems. Builds and evaluates gating networks, selection policies, load‑balancing and sparsity constraints, and associated training and analysis methods to trade off compute and latency against performance, measure expert utilization, and ensure robust, adaptive expert selection.

mixture-of-expertsrouting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.27
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$209K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work investigates the information-theoretic efficiency of routing mechanisms in sparse Mixture-of-Experts (MoE) architectures, aiming to balance model accuracy with communication and computational resource utilization. The gating router is modeled as a stochastic channel, and a discrete mutual information estimator is proposed under a finite expert pool. Empirical posterior distributions \( q(W|S) \) are leveraged to compute \( I(X;T) \) and \( I(S;W) \), with the latter shown to exhibit a monotonic relationship with the generalization gap. The Blahut–Arimoto algorithm is employed to trace the accuracy–rate trade-off curve. Experiments demonstrate that the proposed mutual information estimator effectively tracks the generalization gap and significantly outperforms both the Xu–Raginsky bound and the uniform joint bound, offering a practical analytical tool for resource-aware MoE systems.

communication efficiencyexpert routingfinite expert bank

Existing research on vision-based Mixture-of-Experts (MoE) models predominantly relies on category-level routing statistics, which obscures the actual representational content encoded by individual experts. This work trains sparsely gated convolutional MoE models and advances expert analysis from categorical labels to continuous visual and semantic feature dimensions for the first time. By integrating contrastive learning, neuroscience-inspired tuning analyses, and representational similarity analysis (RSA)—augmented with human semantic judgments from the THINGS dataset to define semantic axes—we demonstrate that experts consistently differentiate along continuous semantic dimensions such as “animate–inanimate.” Despite sparse routing, experts collectively span a broad semantic space. While experts exhibit comparable category discriminability, their feature tuning profiles differ markedly, underscoring the necessity and efficacy of expert-level representational analysis.

expert specialisationMixture-of-Expertsrouting

This work addresses the routing collapse and expert deadlocks that commonly afflict Token-Choice sparse Mixture-of-Experts (MoE) architectures in video diffusion Transformers, which severely limit expert diversity utilization. Starting from a 5-billion-parameter dense model, the authors formulate three principles for converting dense networks to MoE. Through temporal routing analysis of 65 million tokens, they reveal that deadlocked layers follow a U-shaped distribution across the network depth and propose a “functional redundancy” hypothesis to explain this phenomenon. Building on these insights, they integrate expert cloning, zero-initialized gating, auxiliary losses, and enhanced router designs—including linear, MLP, and cross-attention variants—to effectively mitigate bfloat16 precision pitfalls. Their approach alleviates single-expert deadlocks in approximately two-thirds of network layers, endows the model with partial self-recovery capability, delineates the capacity limits of the Token-Choice paradigm, and outlines a three-stage roadmap toward unified vision models and ultimately world models.

Diffusion TransformersRouting CollapseSelective Deadlock

Quadratic Gating Functions in Mixture of Experts: A Statistical Insight

Oct 15, 2024
PA
Pedram Akbarian
🏛️ The University of Texas at Austin | Johns Hopkins University

This work addresses two key limitations in mixture-of-experts (MoE) models: the lack of theoretical connection between MoE routing and self-attention, and the low sample efficiency of linear gating. We propose quadratic gating—replacing conventional linear routing with a quadratic function—and establish, for the first time, its rigorous equivalence to self-attention. Leveraging this equivalence, we derive principled design criteria for optimal quadratic gating and expert functions, leading to a novel high-performance attention mechanism. Theoretically, via statistical learning analysis, we prove that quadratic gating substantially enhances the expressivity and parameter/sample efficiency of expert selection. Empirically, our MoE variant outperforms linear-gating baselines across multiple tasks; the new attention mechanism surpasses state-of-the-art methods—including FlashAttention and Multi-Head Attention—while exhibiting strong alignment between theoretical predictions and empirical results. The framework thus achieves both interpretability and practical efficacy.

Analyzes convergence of MoE models with quadratic gating functionsEstablishes connection between MoE and self-attention mechanismsProposes active-attention mechanism to enhance self-attention performance

Latest Papers

What's happening recently
View more

This study addresses the limited interpretability of expert routing mechanisms in Mixture-of-Experts (MoE) models and the shortcomings of the prevailing "single-domain specialization" assumption by proposing a novel "superimposed specialization" paradigm. To this end, we introduce RouterInterp, a method that integrates sparse autoencoders, feature prediction, and natural language generation to elucidate the intrinsic relationship between sparse features and routing decisions while automatically producing natural language explanations. This approach enables precise mapping from broad domains to fine-grained features. Evaluated on the gpt-oss-20b model, RouterInterp achieves an approximate 65% improvement in detection accuracy over conventional baselines, substantially enhancing the interpretability of MoE architectures.

Expert RoutingInterpretabilityMixture of Experts

This work systematically dissects the multidimensional design space of Mixture-of-Experts (MoE) architectures in large language models, moving beyond conventional generational narratives. It introduces a five-dimensional analytical framework encompassing expert granularity, topology, routing flexibility, load-balancing scope, and execution structure, and constructs a dependency graph to elucidate the coupling mechanisms across four control planes: expert topology, routing, load balancing, and expert parallelism. The framework’s validity is empirically demonstrated through iso-budget pretraining experiments integrating algorithmic innovations—such as Top-k routing, shared and fine-grained experts, and dynamic expert composition—with system-level optimizations including token dispatch, device placement, and all-to-all communication. The study further distills key open challenges for the future development of MoE systems.

expert parallelismexpert topologyload balancing

This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.

dynamic routingexpert capacityMixture-of-Experts

This study investigates the impact of expert granularity on model performance within dense Mixture-of-Experts (MoE) architectures under a fixed computational budget. Theoretically, we demonstrate that dense MoE is equivalent to a standard feed-forward network (FFN) with token-dependent scaling, revealing a non-monotonic effect of fine-grained experts. Empirically, employing SwiGLU activations, Softmax gating, and fixed-width FFNs, we systematically compare varying numbers of activated experts in terms of validation loss and routing behavior. Our findings indicate that activating K=2 experts yields optimal performance, significantly outperforming the baseline, whereas higher granularity paradoxically degrades results. These observations confirm the critical role of soft gating mechanisms under specific architectural configurations.

Dense MoEFeed-Forward NetworkFixed Compute

This study investigates how the routing mechanism of the Mixtral 8x7B-Instruct model influences safety outcomes in response to both benign and harmful prompts. By jointly analyzing expert activation frequencies and router gating gradients—and integrating targeted expert suppression with cross-group expert categorization—the work reveals, for the first time, the deep dependency and distributed nature of safety-related routing decisions. The findings demonstrate that safety-critical experts are broadly dispersed yet concentrated in specific layers; moreover, selectively suppressing experts identified via gradient-based importance significantly reduces restricted responses while inducing fewer side effects, thereby overcoming the limitations of single-metric analyses.

harmful promptslanguage modelsmixture-of-experts

Hot Scholars

CG

Chenjuan Guo

Professor, East China Normal University
Data AnalyticsMachine Learning
DX

Derong Xu

University of Sicence and Technology of China; City University of Hong Kong
Large Language ModelsKnowledge GraphMultimodal Learning
KG

Kun Gai

Senior Director & Researcher, Alibaba Group
Machine LearningComputational Advertising
SS

Sambit Sahu

Capital One
Generative AILLM Pre-trainingInference Optimization
YH

Yao Hu

浙江大学
Machine Learning