Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the storage and deployment bottlenecks caused by parameter redundancy in large Mixture-of-Experts (MoE) models, as well as the irreducible errors in existing pruning-merging methods arising from routing and expert heterogeneity. We propose SLBF, a data-free weight reconstruction framework that establishes the first structural error bounds for pruning and merging. By introducing shared low-rank factorization and post-hoc canonical fixation, SLBF achieves efficient cross-expert compression without requiring original training data while fully preserving routing mechanisms. Evaluated across five MoE architectures ranging from 16B to 122B parameters, SLBF demonstrates lower reconstruction error and faster convergence, comprehensively outperforming three mainstream compression approaches.
📝 Abstract
Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges. We analyze three MoE compression families: expert pruning, expert merging, and weight reconstruction, and derive structural error bounds showing that pruning and merging can incur non-vanishing errors tied to routing and expert heterogeneity. In contrast, weight reconstruction avoids these structural costs by preserving expert structure and routing. Motivated by the analysis, we propose Shared Low-rank Basis Factorization (SLBF), a data-free weight reconstruction method that uses rank-$k$ bases shared among experts, enabling richer cross-expert sharing, faster convergence, and lower reconstruction error. A post-hoc gauge fixing removes redundant parameters at no representational cost. Across five MoE architectures spanning 16B to 122B parameters, SLBF consistently outperforms methods from all three compression families.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Model Compression
Data-free
Large Language Models
Weight Reconstruction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Data-free Compression
Shared Low-rank Basis Factorization
Weight Reconstruction
Gauge Fixing