Score
Designs and implements feature extraction systems that produce representations conditioned on or partitioned by input domains, typically by composing multiple expert extractors with routing or gating mechanisms (mixture-of-experts) to capture domain-specific structure while enabling shared capacity. Engineers and evaluates the extractor architecture, routing policies, and training/regularization procedures to reduce inter-domain representation interference and stabilize latent representations across domains.
This work addresses a critical gap in existing literature by providing the first comprehensive survey of sparse Mixture-of-Experts (MoE) models, systematically integrating their algorithmic foundations, decentralized architectures, and applications in vertical domains. It thoroughly examines core mechanisms such as routing strategies and expert network design, while further extending the discussion to decentralized deployment paradigms and adaptation methods for cross-modal and domain-specific scenarios. By synthesizing recent advances across these dimensions, this survey fills a notable void in the current body of review literature and offers an authoritative reference for researchers and practitioners aiming to develop efficient, scalable large models grounded in sparse MoE principles.
This paper addresses two critical challenges in large language model development: excessive computational overhead and difficulty in modeling heterogeneous, complex data. To tackle these, we present a systematic, up-to-date survey of Mixture-of-Experts (MoE) models. Unlike prior surveys—often outdated or narrowly scoped—we unify and analyze MoE advancements across emerging paradigms including continual learning, meta-learning, and reinforcement learning. We propose a comprehensive framework integrating theoretical analysis (e.g., convergence guarantees), multimodal adaptation (vision and language), and systems-level optimizations (sparse routing, load balancing, distributed training). Furthermore, we introduce a taxonomy of future research directions. Our work establishes the most complete MoE knowledge graph to date, explicitly identifying key bottlenecks and viable technical pathways. It serves as both a methodological foundation and an engineering roadmap for developing efficient, scalable large models.
In high-dimensional sparse Mixture-of-Experts (MoE) models, conventional routers struggle to discern latent token clustering structures, resulting in slow convergence, poor robustness against data corruption, and degraded representation learning. To address this, we propose the Adaptive Clustering (AC) router: it employs a learnable feature-weighting mapping to project tokens into a latent space conducive to expert separation; and introduces, for the first time, an expert-level compactness-aware dynamic feature weighting mechanism—enabling each expert to specialize within semantically coherent subspaces. The method integrates clustering optimization theory, adaptive feature scaling, and token-level routing reparameterization. Evaluated on language modeling and image recognition tasks, the AC router achieves significant improvements in convergence speed, robustness, and overall performance, effectively mitigating routing failure induced by clustering unidentifiability in high-dimensional spaces.
本文通过局部聚合视角分析混合专家模型,探讨路由、稀疏激活和共享专家等设计选择的统计作用,分离出逼近误差、专家学习误差和路由器估计误差。
This work addresses the limitation of existing domain generalization methods that enforce global invariance, thereby overlooking transferable features shared only among subsets of source domains and constraining representational capacity. To overcome this, the authors propose a subset-shared invariance hypothesis and introduce a mixture-of-experts architecture to learn localized invariances: each expert aligns representations within a specific subset of domains, while a routing mechanism dynamically combines expert outputs. The approach jointly optimizes feature representations, selective alignment objectives, and a confidence-balanced routing strategy, further enhanced by a training scheme that encourages diverse expert specialization. Evaluated on the DomainBed benchmark, the method achieves substantial improvements in out-of-domain generalization and demonstrates heightened robustness under increased domain heterogeneity.
This work challenges the prevailing assumption that Mixture-of-Experts (MoE) models achieve domain specialization through sparse routing, introducing the COMMITTEEAUDIT framework to systematically analyze expert-level routing behavior. Through quantitative and qualitative evaluation of multiple representative MoE models on the MMLU benchmark, we uncover the existence of persistent “standing committees”—a small subset of experts that consistently dominate routing weights across domains and layers, regardless of routing budget constraints. These core experts anchor structural and syntactic reasoning, while peripheral experts handle only narrow, domain-specific knowledge. Our findings reveal that the actual degree of specialization in MoE models is substantially lower than commonly assumed and suggest that current load-balancing training objectives may conflict with the model’s intrinsic optimization dynamics.
This work investigates the opaque expert specialization mechanism in Mixture-of-Experts (MoE) models, which limits inference and memory efficiency. By analyzing domain-specific routing patterns and employing an early-decoding framework, the study systematically examines how individual experts contribute to model outputs. Through comprehensive analyses—including routing distribution statistics, cosine similarity of hidden states, comparisons between single-expert and ensemble outputs, and perplexity evaluation—the authors find that a small subset of experts handles over 50% of all requests. Remarkably, outputs from a single dominant expert exhibit high consistency with the full model (cosine similarity up to 0.95), with only a 5% increase in perplexity. These findings suggest that precise expert pruning can substantially enhance inference efficiency without compromising performance, offering a promising avenue for efficient MoE deployment and knowledge localization.
This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.
This work addresses the routing collapse and expert deadlocks that commonly afflict Token-Choice sparse Mixture-of-Experts (MoE) architectures in video diffusion Transformers, which severely limit expert diversity utilization. Starting from a 5-billion-parameter dense model, the authors formulate three principles for converting dense networks to MoE. Through temporal routing analysis of 65 million tokens, they reveal that deadlocked layers follow a U-shaped distribution across the network depth and propose a “functional redundancy” hypothesis to explain this phenomenon. Building on these insights, they integrate expert cloning, zero-initialized gating, auxiliary losses, and enhanced router designs—including linear, MLP, and cross-attention variants—to effectively mitigate bfloat16 precision pitfalls. Their approach alleviates single-expert deadlocks in approximately two-thirds of network layers, endows the model with partial self-recovery capability, delineates the capacity limits of the Token-Choice paradigm, and outlines a three-stage roadmap toward unified vision models and ultimately world models.
This study addresses the limited interpretability of expert routing mechanisms in Mixture-of-Experts (MoE) models and the shortcomings of the prevailing "single-domain specialization" assumption by proposing a novel "superimposed specialization" paradigm. To this end, we introduce RouterInterp, a method that integrates sparse autoencoders, feature prediction, and natural language generation to elucidate the intrinsic relationship between sparse features and routing decisions while automatically producing natural language explanations. This approach enables precise mapping from broad domains to fine-grained features. Evaluated on the gpt-oss-20b model, RouterInterp achieves an approximate 65% improvement in detection accuracy over conventional baselines, substantially enhancing the interpretability of MoE architectures.
该研究通过引入包含决策树、线性支持向量机和二次判别分析的异构专家混合框架,解决了现有可解释MoE模型单一归纳偏置的问题。
This study addresses the substantial computational overhead and inefficiency encountered when adapting multimodal Mixture-of-Experts (MoE) models to specific domains. To this end, it proposes ExpertLens, a data-free framework that exploits the semantic modularity inherent in MoE sparsity. By decoding pretrained router weights, the method precisely identifies domain-relevant experts and applies selective fine-tuning exclusively to their associated parameters. This work highlights the semantic modularity value of sparse architectures. Experimental results demonstrate that updating merely 21.7%–47.0% of the parameters achieves performance on par with full fine-tuning while accelerating training up to fourfold, significantly outperforming mainstream baselines such as LoRA.