A Comprehensive Survey of Mixture-of-Experts: Algorithms, Theory, and Applications

📅 2025-03-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses two critical challenges in large language model development: excessive computational overhead and difficulty in modeling heterogeneous, complex data. To tackle these, we present a systematic, up-to-date survey of Mixture-of-Experts (MoE) models. Unlike prior surveys—often outdated or narrowly scoped—we unify and analyze MoE advancements across emerging paradigms including continual learning, meta-learning, and reinforcement learning. We propose a comprehensive framework integrating theoretical analysis (e.g., convergence guarantees), multimodal adaptation (vision and language), and systems-level optimizations (sparse routing, load balancing, distributed training). Furthermore, we introduce a taxonomy of future research directions. Our work establishes the most complete MoE knowledge graph to date, explicitly identifying key bottlenecks and viable technical pathways. It serves as both a methodological foundation and an engineering roadmap for developing efficient, scalable large models.

Technology Category

Machine Learning: Mixture of Experts (MoE)Natural Language Processing: (Large) Language ModelsComputer Vision: Large Vision Models

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Artificial intelligence (AI) has achieved astonishing successes in many domains, especially with the recent breakthroughs in the development of foundational large models. These large models, leveraging their extensive training data, provide versatile solutions for a wide range of downstream tasks. However, as modern datasets become increasingly diverse and complex, the development of large AI models faces two major challenges: (1) the enormous consumption of computational resources and deployment difficulties, and (2) the difficulty in fitting heterogeneous and complex data, which limits the usability of the models. Mixture of Experts (MoE) models has recently attracted much attention in addressing these challenges, by dynamically selecting and activating the most relevant sub-models to process input data. It has been shown that MoEs can significantly improve model performance and efficiency with fewer resources, particularly excelling in handling large-scale, multimodal data. Given the tremendous potential MoE has demonstrated across various domains, it is urgent to provide a comprehensive summary of recent advancements of MoEs in many important fields. Existing surveys on MoE have their limitations, e.g., being outdated or lacking discussion on certain key areas, and we aim to address these gaps. In this paper, we first introduce the basic design of MoE, including gating functions, expert networks, routing mechanisms, training strategies, and system design. We then explore the algorithm design of MoE in important machine learning paradigms such as continual learning, meta-learning, multi-task learning, and reinforcement learning. Additionally, we summarize theoretical studies aimed at understanding MoE and review its applications in computer vision and natural language processing. Finally, we discuss promising future research directions.
Problem

Research questions and friction points this paper is trying to address.

Addresses computational resource challenges in large AI models
Improves handling of diverse and complex datasets with MoE
Summarizes advancements and applications of Mixture-of-Experts models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic sub-model selection for efficiency
Handling large-scale, multimodal data effectively
Comprehensive MoE advancements across domains
💼 Related Jobs
No related jobs found.
S
Siyuan Mu
College of Science, Sichuan Agricultural University, Ya’an, Sichuan, China
S
Sen Lin
Department of Computer Science, University of Houston, Houston, TX, USA