multi-task mixture of experts

Designs and implements multi-task mixture-of-experts (MoE) models that jointly learn multiple tasks by training a pool of expert subnetworks and gating mechanisms to route inputs and share experts across tasks. Builds and analyzes training procedures, gating/routing strategies, expert capacity and regularization to scale model capacity, reduce task interference, and balance task-specific versus shared representations.

multi-taskmixtureofexperts

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.54
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Mixture of Experts in Large Language Models

Jul 15, 2025
DZ
Danyang Zhang
🏛️ ByteDance Inc | Imperial College London | Purdue University | Heriot-Watt University | Vokram Group | Singapore General Hospital

This paper presents a systematic survey of recent advances in Mixture-of-Experts (MoE) architectures for large language models. Addressing the fundamental trade-off between model capacity scaling and computational efficiency, it investigates key directions: expert gating and dynamic routing mechanisms, hierarchical sparse structure design, meta-learning–enhanced expert collaboration, multimodal/multitask adaptation, and practical deployment challenges. The work proposes a novel MoE effectiveness enhancement framework centered on expert diversity modeling, gating calibration optimization, and improved reliability of inference-time expert aggregation—demonstrating significant gains over both dense models and Bayesian baselines of comparable parameter count. Beyond empirical advances, the study identifies critical bottlenecks—including expert load imbalance, training instability, and hardware inefficiency—and establishes a principled theoretical framework alongside actionable guidelines for designing efficient, scalable MoE-based LLMs. (149 words)

Addressing challenges in expert diversity and inference reliabilityAnalyzing expert gating and routing mechanisms in MoEEnhancing model performance with minimal computational overhead

CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuning

May 23, 2025
JF
Jinyuan Feng
🏛️ Chinese Academy of Sciences | University of Chinese Academy of Sciences | University of Science and Technology Beijing

In parameter-efficient fine-tuning of Mixture-of-Experts (MoE) models, expert functional overlap and suboptimal capacity utilization persist under heterogeneous data distributions. Method: This paper proposes the first contrastive learning framework for sparse-gated MoE fine-tuning, built upon top-k routing. It constructs contrastive objectives between activated and unactivated experts per input, explicitly modeling mutual information differences between inputs and experts to enhance expert modularity and task specificity. Mutual information is approximated and optimized to sharpen expert specialization and improve capacity utilization. Contribution/Results: Experiments demonstrate consistent performance gains across multi-task and standard benchmarks, with measurable improvements in expert specialization—while maintaining identical computational overhead.

Addressing knowledge overlap among experts in heterogeneous datasetsEnhancing MoE capacity and modularization via contrastive representationImproving expert specialization in MoE for parameter-efficient fine-tuning

This work investigates the opaque expert specialization mechanism in Mixture-of-Experts (MoE) models, which limits inference and memory efficiency. By analyzing domain-specific routing patterns and employing an early-decoding framework, the study systematically examines how individual experts contribute to model outputs. Through comprehensive analyses—including routing distribution statistics, cosine similarity of hidden states, comparisons between single-expert and ensemble outputs, and perplexity evaluation—the authors find that a small subset of experts handles over 50% of all requests. Remarkably, outputs from a single dominant expert exhibit high consistency with the full model (cosine similarity up to 0.95), with only a 5% increase in perplexity. These findings suggest that precise expert pruning can substantially enhance inference efficiency without compromising performance, offering a promising avenue for efficient MoE deployment and knowledge localization.

expert specializationinference optimizationMixture of Experts

Theory on Mixture-of-Experts in Continual Learning

Jun 24, 2024
HL
Hongbo Li
🏛️ Singapore University of Technology and Design | University of Houston | The Ohio State University

Existing mixture-of-experts (MoE) models for continual learning lack rigorous theoretical foundations. Method: This paper establishes the first theoretical analysis framework for MoE in continual learning, based on overparameterized linear regression. It derives explicit closed-form expressions for both forgetting error and generalization error, characterizes how expert specialization and dynamic routing jointly mitigate catastrophic forgetting, proves that gating networks must be selectively frozen to ensure convergence, and quantifies the trade-off between the number of experts and convergence iterations. Results: The theory demonstrates that MoE strictly outperforms single-expert models. Extensive experiments on synthetic and real-world benchmarks with deep neural networks empirically validate the theoretical predictions, confirming the efficacy of expert specialization, optimal gating freezing schedules, and the expert-count–convergence trade-off. This work provides the first interpretable, theoretically grounded foundation for MoE-based continual learning.

Impact of expert count on convergence and performance.Mitigating catastrophic forgetting via expert diversification.Theoretical analysis of Mixture-of-Experts in Continual Learning.

A Closer Look into Mixture-of-Experts in Large Language Models

Jun 26, 2024
KM
Ka Man Lo
🏛️ University of Macau | University of Edinburgh | Tsinghua University | INF Technology | HKUST

The internal mechanisms and modular nature of Mixture-of-Experts (MoE) large language models remain poorly understood, particularly regarding expert granularity, routing behavior, and layer-wise expert diversity. Method: We conduct attribution analysis, expert activation visualization, output norm statistics, and controlled experiments across three representative MoE architectures—Mixtral, GLaM, and DeepSpeed-MoE. Contribution/Results: We empirically establish that individual neurons function as fine-grained experts; routers exhibit strong preference for high-norm experts; and expert diversity generally increases with network depth—except in the final layer. Based on these findings, we formulate a hierarchical evolution law of expert diversity and provide actionable guidelines for router design and expert allocation. Our work formally validates the modular architecture of MoE models, identifies anomalous behavior in the top layer, and has directly informed routing strategy improvements across multiple research teams. The open-sourced code has garnered significant community attention.

Exploring parametric and behavioral features of MoE modelsInvestigating expert diversity and router selection mechanismsUnderstanding inner workings of MoE-based large language models

Latest Papers

What's happening recently
View more

This work proposes the DS-MoE framework to address the memory bottlenecks in deploying Mixture-of-Experts (MoE) models and the limitation of conventional Top-k routing in neglecting expert dependencies. The framework reveals the functional duality of expert combinations and formulates the selection objective as a difference-of-submodular function, thereby decoupling redundancy from synergy effects. Furthermore, based on second-order Taylor expansion analysis, it designs a majorization-minimization algorithm with monotonicity guarantees to achieve dependency-aware extraction of compact expert subsets. Experimental results demonstrate that the proposed method effectively preserves critical expert combinations and significantly outperforms existing baselines.

Expert SelectionInter-expert DependenciesMemory Efficiency

This study addresses the lack of systematic investigation into the role of Mixture-of-Experts (MoE) in multimodal learning. It presents the first integrative analytical framework bridging MoE and multimodal learning, examining its applications through three complementary lenses: as an efficient computational engine, a representation learner, and an adapter. The work systematically reviews how MoE enhances computational scalability, cross-modal alignment, and modeling under imperfect data conditions. Drawing on a comprehensive literature survey, it synthesizes key technical aspects—including routing mechanisms, expert selection, representation alignment, and handling of missing modalities—into a unified theoretical framework. The paper further identifies critical research gaps, such as interpretable routing, inter-expert communication, adaptive modality fusion, and continual learning, thereby charting a path toward building interpretable and sustainable multimodal MoE systems.

missing modalityMixture-of-Expertsmodality imbalance

This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.

dynamic routingexpert capacityMixture-of-Experts

Hot Scholars

WJ

Weisen Jiang

CUHK, HKUST
large language modelsdeep learningmeta-learning
BL

Baijiong Lin

Ph.D. Student, The Hong Kong University of Science and Technology (Guangzhou)
RLVRLLM Post-TrainingMulti-Task Learning
FY

Feiyang Ye

University of Technology Sydney, Ph.D student
Multi-Task Learning