OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit

📅 2026-09-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory bottleneck caused by expert redundancy in large Mixture-of-Experts (MoE) models, as well as the high search costs and neglect of dynamic dependencies in existing pruning methods, by proposing a training-free compression framework. Methodologically, it introduces the first expert pruning paradigm based on orthogonal matching pursuit, reformulating pruning as sparse signal reconstruction. A water-filling strategy is incorporated to balance reconstruction quality with routing stability, alongside an energy-prediction-driven mechanism for cross-layer allocation and adaptive inference. Evaluated on models such as Qwen, the proposed framework achieves compression ratios of 25%–50% while retaining 93.3% of the original performance. Furthermore, it accelerates the pruning search process by 33× and improves inference speed by 1.55×.
📝 Abstract
Mixture-of-Experts (MoE) models enable efficient scaling of large language models but face critical deployment challenges due to massive memory requirements. Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts. To address these challenges, we present OMP-MoE, a novel training-free compression framework for reducing expert redundancy in MoE-based LLMs. Based on observations of expert contribution patterns, we reformulate the pruning problem as a sparse signal reconstruction task solved through Orthogonal Matching Pursuit. Specifically, our method first treats individual expert contributions as dictionary atoms and selects experts that greedily minimize reconstruction error with linear computational complexity. Then, we optimize cross-layer expert allocation through a water-filling strategy that accounts for both reconstruction quality and routing stability. Finally, we introduce OMP-MoE{\dag}, an adaptive inference mechanism that dynamically adjusts expert activation based on energy prediction. Comprehensive experiments on Qwen, DeepSeek-V2, GPT-OSS, and Mixtral MoE demonstrate consistent improvements over existing methods at 25-50% pruning ratios. For Qwen3-30B-A3B at 50% compression, we retain 93.3% of original performance, achieving 33$\times$ faster search and 1.55$\times$ inference speedup. Codes will be available after acceptance.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Large Language Models
Expert Pruning
Model Compression
Memory Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Orthogonal Matching Pursuit
Training-free Pruning
Sparse Signal Reconstruction
Adaptive Inference
💼 Related Jobs
No related jobs found.
Dezhi Li
Dezhi Li
University of Waterloo
Prognostics and Health ManagementSignal ProcessingMachine LearningArtificial Intelligence
L
Lujun Li
The Hong Kong University of Science and Technology
Q
Qiyuan Zhu
The Hong Kong University of Science and Technology
Hao Gu
Hao Gu
Sun Yat-Sen University
Planetary aeronomyAtmospheric escapeSpace physics
B
Bei Liu
The Hong Kong University of Science and Technology
Sirui Han
Sirui Han
The Hong Kong University of Science and Technology
Large Language ModelInterdisciplinary Artificial Intelligence
Y
Yike Guo
The Hong Kong University of Science and Technology