Score
Designs and implements mechanisms that route inputs, internal activations, or contextual signals to a subset of specialist model components (‘experts’) — covering conditional, class- or task-conditioned, dynamic, state-gated, frequency- or residual-guided, scale-adaptive, sparse vs dense, bi-modal or autoencoder-based routing and bridge/gating variants — so different experts handle different subproblems. Builds and evaluates gating networks, selection policies, load‑balancing and sparsity constraints, and associated training and analysis methods to trade off compute and latency against performance, measure expert utilization, and ensure robust, adaptive expert selection.
This work addresses a critical gap in existing literature by providing the first comprehensive survey of sparse Mixture-of-Experts (MoE) models, systematically integrating their algorithmic foundations, decentralized architectures, and applications in vertical domains. It thoroughly examines core mechanisms such as routing strategies and expert network design, while further extending the discussion to decentralized deployment paradigms and adaptation methods for cross-modal and domain-specific scenarios. By synthesizing recent advances across these dimensions, this survey fills a notable void in the current body of review literature and offers an authoritative reference for researchers and practitioners aiming to develop efficient, scalable large models grounded in sparse MoE principles.
本文通过局部聚合视角分析混合专家模型,探讨路由、稀疏激活和共享专家等设计选择的统计作用,分离出逼近误差、专家学习误差和路由器估计误差。
This work investigates the information-theoretic efficiency of routing mechanisms in sparse Mixture-of-Experts (MoE) architectures, aiming to balance model accuracy with communication and computational resource utilization. The gating router is modeled as a stochastic channel, and a discrete mutual information estimator is proposed under a finite expert pool. Empirical posterior distributions \( q(W|S) \) are leveraged to compute \( I(X;T) \) and \( I(S;W) \), with the latter shown to exhibit a monotonic relationship with the generalization gap. The Blahut–Arimoto algorithm is employed to trace the accuracy–rate trade-off curve. Experiments demonstrate that the proposed mutual information estimator effectively tracks the generalization gap and significantly outperforms both the Xu–Raginsky bound and the uniform joint bound, offering a practical analytical tool for resource-aware MoE systems.
Existing research on vision-based Mixture-of-Experts (MoE) models predominantly relies on category-level routing statistics, which obscures the actual representational content encoded by individual experts. This work trains sparsely gated convolutional MoE models and advances expert analysis from categorical labels to continuous visual and semantic feature dimensions for the first time. By integrating contrastive learning, neuroscience-inspired tuning analyses, and representational similarity analysis (RSA)—augmented with human semantic judgments from the THINGS dataset to define semantic axes—we demonstrate that experts consistently differentiate along continuous semantic dimensions such as “animate–inanimate.” Despite sparse routing, experts collectively span a broad semantic space. While experts exhibit comparable category discriminability, their feature tuning profiles differ markedly, underscoring the necessity and efficacy of expert-level representational analysis.
This work addresses the routing collapse and expert deadlocks that commonly afflict Token-Choice sparse Mixture-of-Experts (MoE) architectures in video diffusion Transformers, which severely limit expert diversity utilization. Starting from a 5-billion-parameter dense model, the authors formulate three principles for converting dense networks to MoE. Through temporal routing analysis of 65 million tokens, they reveal that deadlocked layers follow a U-shaped distribution across the network depth and propose a “functional redundancy” hypothesis to explain this phenomenon. Building on these insights, they integrate expert cloning, zero-initialized gating, auxiliary losses, and enhanced router designs—including linear, MLP, and cross-attention variants—to effectively mitigate bfloat16 precision pitfalls. Their approach alleviates single-expert deadlocks in approximately two-thirds of network layers, endows the model with partial self-recovery capability, delineates the capacity limits of the Token-Choice paradigm, and outlines a three-stage roadmap toward unified vision models and ultimately world models.
This work addresses two key limitations in mixture-of-experts (MoE) models: the lack of theoretical connection between MoE routing and self-attention, and the low sample efficiency of linear gating. We propose quadratic gating—replacing conventional linear routing with a quadratic function—and establish, for the first time, its rigorous equivalence to self-attention. Leveraging this equivalence, we derive principled design criteria for optimal quadratic gating and expert functions, leading to a novel high-performance attention mechanism. Theoretically, via statistical learning analysis, we prove that quadratic gating substantially enhances the expressivity and parameter/sample efficiency of expert selection. Empirically, our MoE variant outperforms linear-gating baselines across multiple tasks; the new attention mechanism surpasses state-of-the-art methods—including FlashAttention and Multi-Head Attention—while exhibiting strong alignment between theoretical predictions and empirical results. The framework thus achieves both interpretability and practical efficacy.
This study addresses the limited interpretability of expert routing mechanisms in Mixture-of-Experts (MoE) models and the shortcomings of the prevailing "single-domain specialization" assumption by proposing a novel "superimposed specialization" paradigm. To this end, we introduce RouterInterp, a method that integrates sparse autoencoders, feature prediction, and natural language generation to elucidate the intrinsic relationship between sparse features and routing decisions while automatically producing natural language explanations. This approach enables precise mapping from broad domains to fine-grained features. Evaluated on the gpt-oss-20b model, RouterInterp achieves an approximate 65% improvement in detection accuracy over conventional baselines, substantially enhancing the interpretability of MoE architectures.
This work systematically dissects the multidimensional design space of Mixture-of-Experts (MoE) architectures in large language models, moving beyond conventional generational narratives. It introduces a five-dimensional analytical framework encompassing expert granularity, topology, routing flexibility, load-balancing scope, and execution structure, and constructs a dependency graph to elucidate the coupling mechanisms across four control planes: expert topology, routing, load balancing, and expert parallelism. The framework’s validity is empirically demonstrated through iso-budget pretraining experiments integrating algorithmic innovations—such as Top-k routing, shared and fine-grained experts, and dynamic expert composition—with system-level optimizations including token dispatch, device placement, and all-to-all communication. The study further distills key open challenges for the future development of MoE systems.
This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.
This study investigates the impact of expert granularity on model performance within dense Mixture-of-Experts (MoE) architectures under a fixed computational budget. Theoretically, we demonstrate that dense MoE is equivalent to a standard feed-forward network (FFN) with token-dependent scaling, revealing a non-monotonic effect of fine-grained experts. Empirically, employing SwiGLU activations, Softmax gating, and fixed-width FFNs, we systematically compare varying numbers of activated experts in terms of validation loss and routing behavior. Our findings indicate that activating K=2 experts yields optimal performance, significantly outperforming the baseline, whereas higher granularity paradoxically degrades results. These observations confirm the critical role of soft gating mechanisms under specific architectural configurations.
This study investigates how the routing mechanism of the Mixtral 8x7B-Instruct model influences safety outcomes in response to both benign and harmful prompts. By jointly analyzing expert activation frequencies and router gating gradients—and integrating targeted expert suppression with cross-group expert categorization—the work reveals, for the first time, the deep dependency and distributed nature of safety-related routing decisions. The findings demonstrate that safety-critical experts are broadly dispersed yet concentrated in specific layers; moreover, selectively suppressing experts identified via gradient-based importance significantly reduces restricted responses while inducing fewer side effects, thereby overcoming the limitations of single-metric analyses.