Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models

📅 2026-08-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of efficiently compressing Mixture-of-Experts (MoE) models, which typically require loading all expert parameters. The authors propose a one-shot expert pruning method based on lightweight fine-tuning—such as router-specific LoRA or IA³—that induces changes in router weights. By measuring the ℓ² norm of these weight changes to assess expert sensitivity, they rank and prune the least sensitive experts. This study is the first to demonstrate that router sensitivity serves as an effective pruning signal, enabling near-linear accuracy degradation rather than catastrophic collapse under high compression ratios, with notable transferability across models. On Mixtral-8×7B, pruning 50% of experts yields a 28.76% score on MMLU-Pro, alongside 49% memory reduction and 37% lower latency; on Qwen1.5-MoE, it maintains a 49.7% average accuracy on mathematical tasks, substantially outperforming random or magnitude-based pruning.
📝 Abstract
Mixture-of-Experts (MoE) models decouple total parameters from per-token compute, but deployment still requires storing every expert. Recent theory shows that pruning experts with the smallest router-norm changes during fine-tuning can preserve accuracy, but assumes full fine-tuning. We test whether lightweight adaptation can recover this signal. We briefly fine-tune with a parameter-efficient adapter, rank experts by the induced $\ell_2$ router change, and prune the least-changed experts in one shot. On Mixtral-8$\times$7B-Instruct (44.83% MMLU-Pro), router-only LoRA trains 0.002% of parameters and outperforms all-module LoRA at matched rank with half the experts removed (27.54% vs. 24.42%); signal quality declines as adaptation spreads to attention and expert weights. Accuracy improves monotonically with LoRA rank, reaching 28.76%. IA3, which leaves router weights frozen, matches direct router adaptation, whereas unconstrained additive adapters degrade the signal. Router-guided MMLU-Pro accuracy decays quasi-linearly rather than collapsing, remains nearly 1.8 times that of magnitude-based or random pruning at maximal compression, and reduces memory by 49% and per-token latency by 37%. At 25% compression, retention is competitive with methods using full activation statistics. The criterion also transfers to Qwen1.5-MoE fine-tuned for mathematics, retaining 49.7% mean accuracy over eleven benchmarks with half the experts removed while random pruning falls to single digits. Router sensitivity under lightweight fine-tuning therefore makes provably motivated expert pruning practical at scale.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
expert pruning
router sensitivity
lightweight fine-tuning
model compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
router sensitivity
lightweight fine-tuning
expert pruning
parameter-efficient adaptation
💼 Related Jobs
No related jobs found.
A
Ali Janati
Data Science Institute, Columbia University, New York, NY, USA
K
Kaoutar El Maghraoui
Department of Computer Science, Columbia University, New York, NY, USA
X
Xinyi Luo
Department of Computer Science, Columbia University, New York, NY, USA
W
Wenyuan Shen
Department of Computer Science, Columbia University, New York, NY, USA
O
Owen Zou
Department of Computer Science, Columbia University, New York, NY, USA
Y
Yankai Mao
Department of Computer Science, Columbia University, New York, NY, USA