Redundancy Meets Synergy: Dependency-aware Expert Selection for MoE via Submodular Optimization

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes the DS-MoE framework to address the memory bottlenecks in deploying Mixture-of-Experts (MoE) models and the limitation of conventional Top-k routing in neglecting expert dependencies. The framework reveals the functional duality of expert combinations and formulates the selection objective as a difference-of-submodular function, thereby decoupling redundancy from synergy effects. Furthermore, based on second-order Taylor expansion analysis, it designs a majorization-minimization algorithm with monotonicity guarantees to achieve dependency-aware extraction of compact expert subsets. Experimental results demonstrate that the proposed method effectively preserves critical expert combinations and significantly outperforms existing baselines.
📝 Abstract
While Mixture-of-Experts (MoE) models effectively scale model capacity through sparse activation, their deployment is often bottlenecked by prohibitive memory requirements. Extracting a compact subset of experts presents a promising solution. However, existing expert selection heuristics predominantly rely on Top-k ranking, which isolates the evaluation of individual experts and ignores the intricate inter-expert dependencies introduced by the MoE gating network. In this paper, we propose DS-MoE, a theoretically grounded framework that redefines expert selection via difference-of-submodular (DS) optimization. By analyzing the second-order Taylor expansion of the loss degradation, we reveal functional duality within expert combinations: redundancy (where experts encode overlapping representations) and synergy (where experts provide complementary error cancellation). To navigate this duality, we mathematically decouple redundancy reduction from synergy maximization by formulating the selection objective as a DS function. Furthermore, we devise a tailored majorization-minimization (MM) algorithm with provable monotonicity guarantees to efficiently identify the optimal expert subset. Extensive experiments demonstrate that DS-MoE effectively preserves indispensable expert combinations, achieving superior performance compared to the state-of-the-art baselines.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Expert Selection
Model Compression
Inter-expert Dependencies
Memory Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Submodular Optimization
Expert Selection
Redundancy and Synergy
Majorization-Minimization
💼 Related Jobs
No related jobs found.
Z
Zheng Lin
Interdisciplinary Centre for Security, Reliability and Trust, University of Luxembourg, Luxembourg
S
Shaoke Fang
Department of Computer Science, Peking University, China
Yuxin Zhang
Yuxin Zhang
Fudan University
Distributed Machine LearningEdge AISatellite Internet
J
Jinfeng Xu
Department of Electrical and Computer Engineering, The University of British Columbia, Canada
Z
Zihan Fang
Department of Computer Science, City University of Hong Kong, Hong Kong, China
Zhe Chen
Zhe Chen
Fudan University
Satellite InternetEdge AIWireless Sensing
Wei Ni
Wei Ni
FIEEE, AAIA Fellow, Senior Principal Scientist & Conjoint Professor, CSIRO/UNSW
6G security and privacyconnected and trusted intelligenceapplied AI/ML
J
Jun Luo
College of Computing and Data Science, Nanyang Technological University, Singapore
Symeon Chatzinotas
Symeon Chatzinotas
Full Professor | IEEE Fellow | SIGCOM Head, SnT, University of Luxembourg
Wireless CommunicationsNon-Terrestrial NetworksInternet of Things6GQuantum Communications