Score
Designs, implements, and evaluates sparse mixture-of-experts architectures and their components — learned gating and routing functions, expert selection and hierarchical gating, cross-expert selective attention, and mechanisms for ensembling or averaging expert outputs (including prompt-based or multimodal experts). Builds training, finetuning, pruning, compression-aware and model-averaging procedures and experimental pipelines to assemble, compress, or combine experts and to trade off model capacity, active computation, and accuracy for multi-token or multi-scale outputs.
This work addresses a critical gap in existing literature by providing the first comprehensive survey of sparse Mixture-of-Experts (MoE) models, systematically integrating their algorithmic foundations, decentralized architectures, and applications in vertical domains. It thoroughly examines core mechanisms such as routing strategies and expert network design, while further extending the discussion to decentralized deployment paradigms and adaptation methods for cross-modal and domain-specific scenarios. By synthesizing recent advances across these dimensions, this survey fills a notable void in the current body of review literature and offers an authoritative reference for researchers and practitioners aiming to develop efficient, scalable large models grounded in sparse MoE principles.
This paper addresses two critical challenges in large language model development: excessive computational overhead and difficulty in modeling heterogeneous, complex data. To tackle these, we present a systematic, up-to-date survey of Mixture-of-Experts (MoE) models. Unlike prior surveys—often outdated or narrowly scoped—we unify and analyze MoE advancements across emerging paradigms including continual learning, meta-learning, and reinforcement learning. We propose a comprehensive framework integrating theoretical analysis (e.g., convergence guarantees), multimodal adaptation (vision and language), and systems-level optimizations (sparse routing, load balancing, distributed training). Furthermore, we introduce a taxonomy of future research directions. Our work establishes the most complete MoE knowledge graph to date, explicitly identifying key bottlenecks and viable technical pathways. It serves as both a methodological foundation and an engineering roadmap for developing efficient, scalable large models.
本文通过局部聚合视角分析混合专家模型,探讨路由、稀疏激活和共享专家等设计选择的统计作用,分离出逼近误差、专家学习误差和路由器估计误差。
This work systematically dissects the multidimensional design space of Mixture-of-Experts (MoE) architectures in large language models, moving beyond conventional generational narratives. It introduces a five-dimensional analytical framework encompassing expert granularity, topology, routing flexibility, load-balancing scope, and execution structure, and constructs a dependency graph to elucidate the coupling mechanisms across four control planes: expert topology, routing, load balancing, and expert parallelism. The framework’s validity is empirically demonstrated through iso-budget pretraining experiments integrating algorithmic innovations—such as Top-k routing, shared and fine-grained experts, and dynamic expert composition—with system-level optimizations including token dispatch, device placement, and all-to-all communication. The study further distills key open challenges for the future development of MoE systems.
This work addresses two key limitations in mixture-of-experts (MoE) models: the lack of theoretical connection between MoE routing and self-attention, and the low sample efficiency of linear gating. We propose quadratic gating—replacing conventional linear routing with a quadratic function—and establish, for the first time, its rigorous equivalence to self-attention. Leveraging this equivalence, we derive principled design criteria for optimal quadratic gating and expert functions, leading to a novel high-performance attention mechanism. Theoretically, via statistical learning analysis, we prove that quadratic gating substantially enhances the expressivity and parameter/sample efficiency of expert selection. Empirically, our MoE variant outperforms linear-gating baselines across multiple tasks; the new attention mechanism surpasses state-of-the-art methods—including FlashAttention and Multi-Head Attention—while exhibiting strong alignment between theoretical predictions and empirical results. The framework thus achieves both interpretability and practical efficacy.
This work systematically investigates the interplay among key design dimensions in Mixture-of-Experts (MoE) architectures—such as the number of experts, expert granularity, heterogeneity, shared experts, and load balancing—through over 2,000 large-scale pretraining experiments. The study reveals that the number of experts and their granularity are the dominant factors governing model performance, while other design choices exert comparatively limited influence. Notably, increasing the total MoE parameters consistently enhances performance across all active parameter budgets, and the optimal expert size is determined solely by the number of active parameters. Furthermore, the effectiveness of dropless routing is empirically validated, demonstrating consistent performance gains.
This work investigates the opaque expert specialization mechanism in Mixture-of-Experts (MoE) models, which limits inference and memory efficiency. By analyzing domain-specific routing patterns and employing an early-decoding framework, the study systematically examines how individual experts contribute to model outputs. Through comprehensive analyses—including routing distribution statistics, cosine similarity of hidden states, comparisons between single-expert and ensemble outputs, and perplexity evaluation—the authors find that a small subset of experts handles over 50% of all requests. Remarkably, outputs from a single dominant expert exhibit high consistency with the full model (cosine similarity up to 0.95), with only a 5% increase in perplexity. These findings suggest that precise expert pruning can substantially enhance inference efficiency without compromising performance, offering a promising avenue for efficient MoE deployment and knowledge localization.
This work addresses the storage and memory bottlenecks imposed by the large parameter counts of Mixture-of-Experts (MoE) models in edge-device deployment by proposing DECO, a sparse MoE architecture. DECO integrates differentiable ReLU-based routing, learnable expert-level scaling factors, NormSiLU activation functions, and non-gated MLP experts to achieve performance on par with dense Transformers while activating only 20% of the experts. Compared to existing MoE baselines, DECO substantially enhances sparsity stability and computational efficiency, delivering a 3× inference speedup over dense models on real hardware. The design thus achieves an effective balance among high performance, low computational overhead, and minimal memory footprint.
This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.
研究通过动态专家剪枝方法在细粒度MoE架构中减少冗余专家选择,保留约2/3专家即可保持98.8%性能,提高推理效率。
This work addresses the routing collapse and expert deadlocks that commonly afflict Token-Choice sparse Mixture-of-Experts (MoE) architectures in video diffusion Transformers, which severely limit expert diversity utilization. Starting from a 5-billion-parameter dense model, the authors formulate three principles for converting dense networks to MoE. Through temporal routing analysis of 65 million tokens, they reveal that deadlocked layers follow a U-shaped distribution across the network depth and propose a “functional redundancy” hypothesis to explain this phenomenon. Building on these insights, they integrate expert cloning, zero-initialized gating, auxiliary losses, and enhanced router designs—including linear, MLP, and cross-attention variants—to effectively mitigate bfloat16 precision pitfalls. Their approach alleviates single-expert deadlocks in approximately two-thirds of network layers, endows the model with partial self-recovery capability, delineates the capacity limits of the Token-Choice paradigm, and outlines a three-stage roadmap toward unified vision models and ultimately world models.
为了解决Mixture-of-Experts中参与度、执行成本和内存成本难以独立控制的问题,IntBMoE通过结合块级条件与专家组合的方法实现了全参与的同时保持了低计算和存储成本。