Score
Designs and implements multi-task mixture-of-experts (MoE) models that jointly learn multiple tasks by training a pool of expert subnetworks and gating mechanisms to route inputs and share experts across tasks. Builds and analyzes training procedures, gating/routing strategies, expert capacity and regularization to scale model capacity, reduce task interference, and balance task-specific versus shared representations.
This paper addresses two critical challenges in large language model development: excessive computational overhead and difficulty in modeling heterogeneous, complex data. To tackle these, we present a systematic, up-to-date survey of Mixture-of-Experts (MoE) models. Unlike prior surveys—often outdated or narrowly scoped—we unify and analyze MoE advancements across emerging paradigms including continual learning, meta-learning, and reinforcement learning. We propose a comprehensive framework integrating theoretical analysis (e.g., convergence guarantees), multimodal adaptation (vision and language), and systems-level optimizations (sparse routing, load balancing, distributed training). Furthermore, we introduce a taxonomy of future research directions. Our work establishes the most complete MoE knowledge graph to date, explicitly identifying key bottlenecks and viable technical pathways. It serves as both a methodological foundation and an engineering roadmap for developing efficient, scalable large models.
This paper presents a systematic survey of recent advances in Mixture-of-Experts (MoE) architectures for large language models. Addressing the fundamental trade-off between model capacity scaling and computational efficiency, it investigates key directions: expert gating and dynamic routing mechanisms, hierarchical sparse structure design, meta-learning–enhanced expert collaboration, multimodal/multitask adaptation, and practical deployment challenges. The work proposes a novel MoE effectiveness enhancement framework centered on expert diversity modeling, gating calibration optimization, and improved reliability of inference-time expert aggregation—demonstrating significant gains over both dense models and Bayesian baselines of comparable parameter count. Beyond empirical advances, the study identifies critical bottlenecks—including expert load imbalance, training instability, and hardware inefficiency—and establishes a principled theoretical framework alongside actionable guidelines for designing efficient, scalable MoE-based LLMs. (149 words)
In parameter-efficient fine-tuning of Mixture-of-Experts (MoE) models, expert functional overlap and suboptimal capacity utilization persist under heterogeneous data distributions. Method: This paper proposes the first contrastive learning framework for sparse-gated MoE fine-tuning, built upon top-k routing. It constructs contrastive objectives between activated and unactivated experts per input, explicitly modeling mutual information differences between inputs and experts to enhance expert modularity and task specificity. Mutual information is approximated and optimized to sharpen expert specialization and improve capacity utilization. Contribution/Results: Experiments demonstrate consistent performance gains across multi-task and standard benchmarks, with measurable improvements in expert specialization—while maintaining identical computational overhead.
This work investigates the opaque expert specialization mechanism in Mixture-of-Experts (MoE) models, which limits inference and memory efficiency. By analyzing domain-specific routing patterns and employing an early-decoding framework, the study systematically examines how individual experts contribute to model outputs. Through comprehensive analyses—including routing distribution statistics, cosine similarity of hidden states, comparisons between single-expert and ensemble outputs, and perplexity evaluation—the authors find that a small subset of experts handles over 50% of all requests. Remarkably, outputs from a single dominant expert exhibit high consistency with the full model (cosine similarity up to 0.95), with only a 5% increase in perplexity. These findings suggest that precise expert pruning can substantially enhance inference efficiency without compromising performance, offering a promising avenue for efficient MoE deployment and knowledge localization.
Existing mixture-of-experts (MoE) models for continual learning lack rigorous theoretical foundations. Method: This paper establishes the first theoretical analysis framework for MoE in continual learning, based on overparameterized linear regression. It derives explicit closed-form expressions for both forgetting error and generalization error, characterizes how expert specialization and dynamic routing jointly mitigate catastrophic forgetting, proves that gating networks must be selectively frozen to ensure convergence, and quantifies the trade-off between the number of experts and convergence iterations. Results: The theory demonstrates that MoE strictly outperforms single-expert models. Extensive experiments on synthetic and real-world benchmarks with deep neural networks empirically validate the theoretical predictions, confirming the efficacy of expert specialization, optimal gating freezing schedules, and the expert-count–convergence trade-off. This work provides the first interpretable, theoretically grounded foundation for MoE-based continual learning.
The internal mechanisms and modular nature of Mixture-of-Experts (MoE) large language models remain poorly understood, particularly regarding expert granularity, routing behavior, and layer-wise expert diversity. Method: We conduct attribution analysis, expert activation visualization, output norm statistics, and controlled experiments across three representative MoE architectures—Mixtral, GLaM, and DeepSpeed-MoE. Contribution/Results: We empirically establish that individual neurons function as fine-grained experts; routers exhibit strong preference for high-norm experts; and expert diversity generally increases with network depth—except in the final layer. Based on these findings, we formulate a hierarchical evolution law of expert diversity and provide actionable guidelines for router design and expert allocation. Our work formally validates the modular architecture of MoE models, identifies anomalous behavior in the top layer, and has directly informed routing strategy improvements across multiple research teams. The open-sourced code has garnered significant community attention.
本文提出MetaNet,通过预测每层的专家保留阈值和路由偏置来动态分配专家,解决了MoE模型中固定专家数量导致的效率问题。
本文通过局部聚合视角分析混合专家模型,探讨路由、稀疏激活和共享专家等设计选择的统计作用,分离出逼近误差、专家学习误差和路由器估计误差。
This work proposes the DS-MoE framework to address the memory bottlenecks in deploying Mixture-of-Experts (MoE) models and the limitation of conventional Top-k routing in neglecting expert dependencies. The framework reveals the functional duality of expert combinations and formulates the selection objective as a difference-of-submodular function, thereby decoupling redundancy from synergy effects. Furthermore, based on second-order Taylor expansion analysis, it designs a majorization-minimization algorithm with monotonicity guarantees to achieve dependency-aware extraction of compact expert subsets. Experimental results demonstrate that the proposed method effectively preserves critical expert combinations and significantly outperforms existing baselines.
This study addresses the lack of systematic investigation into the role of Mixture-of-Experts (MoE) in multimodal learning. It presents the first integrative analytical framework bridging MoE and multimodal learning, examining its applications through three complementary lenses: as an efficient computational engine, a representation learner, and an adapter. The work systematically reviews how MoE enhances computational scalability, cross-modal alignment, and modeling under imperfect data conditions. Drawing on a comprehensive literature survey, it synthesizes key technical aspects—including routing mechanisms, expert selection, representation alignment, and handling of missing modalities—into a unified theoretical framework. The paper further identifies critical research gaps, such as interpretable routing, inter-expert communication, adaptive modality fusion, and continual learning, thereby charting a path toward building interpretable and sustainable multimodal MoE systems.
This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.