🤖 AI Summary
This study addresses the substantial memory overhead of deploying Mixture-of-Experts (MoE) models and the limitations of existing pruning methods, including poor alignment, high computational cost, and neglect of routing redundancy. We propose an efficient MoE pruning framework that achieves structured pruning by introducing learnable router biases and diversity regularization to precisely identify critical experts. Furthermore, an affine transformation-based expert approximation mechanism is designed to effectively compensate for accuracy degradation caused by pruning. Experimental results demonstrate that the proposed method successfully removes 25%–50% of experts across multiple large language models while consistently outperforming state-of-the-art algorithms on nine zero-shot benchmarks, achieving a favorable balance between model compression and reasoning capability.
📝 Abstract
Mixture-of-Experts (MoE) models enable parameter scaling with limited per-token computation by activating only a small subset of experts for each token, but deploying them still requires loading the complete expert pool into memory. Structured expert pruning can effectively reduce the memory usage by removing experts. However, existing pruning methods either use expert ranking criteria that are not well aligned with model performance or rely on effective expert subset searching that is computationally expensive. Moreover, these methods typically overlook the routing-behavior redundancy among the retained experts. In this paper, we propose MoE Pruning via Router Bias Learning and Expert Approximation (MoRA), a framework for structured MoE expert pruning. We introduce a learnable router bias for each expert and optimize these biases by minimizing the language-modeling loss and a routing-diversity regularizer. The learned router biases sharpen the routing probability distributions to identify experts critical to model performance while encouraging the selection of experts with diverse routing preferences. In addition, we introduce an expert approximation mechanism as a post-pruning enhancement. It leverages the remaining experts to approximate the outputs of pruned experts by affine transformation, further improving the performance of the pruned model. We evaluate MoRA on Qwen3-30B-A3B, DeepSeek-V2-Lite, and Moonlight-16B-A3B, removing 25\% and 50\% of the routed experts in each MoE layer. Extensive experiments on nine zero-shot benchmarks show that MoRA outperforms state-of-the-art pruning algorithms. Our code will be released.