Score
Designs, implements, and evaluates pruning and compression methods that remove redundant network weights, nodes, expert submodules, or circuit elements while explicitly optimizing hardware costs (area, power, latency, and cross-device communication) and preserving task performance. Builds cost and coverage metrics, task-aware or adaptive prior selection and transfer mechanisms, and importance- or coverage-based expert selection algorithms to decide which components to prune and to predict the resulting execution overhead and accuracy trade-offs.
研究解决了过分散路由下专家修剪导致模型性能下降的问题,提出MESA方法以最小化最坏情况下的领域退化。
Sparse Mixture-of-Experts (MoE) models offer computational efficiency but suffer from functional subspace collapse during expert compression—especially in expert merging—where the router loses input-dependent control over expert selection, inducing irreversible errors. This work is the first to identify and formalize this deficiency. We propose Router-Weighted Expert Pruning (RWEP), a one-shot, fine-tuning-free compression method that jointly leverages gating weights and activation norms to dynamically assess expert importance. We theoretically prove that pruning preserves generative capability more faithfully than merging. Empirical evaluation on models ranging from 20B to 1T parameters shows that 50% expert pruning yields near-lossless performance: Qwen3-Coder-480B and Kimi-K2 achieve competitive code generation and tool-use accuracy, significantly outperforming both expert merging and alternative pruning baselines.
This study addresses the substantial memory overhead of deploying Mixture-of-Experts (MoE) models and the limitations of existing pruning methods, including poor alignment, high computational cost, and neglect of routing redundancy. We propose an efficient MoE pruning framework that achieves structured pruning by introducing learnable router biases and diversity regularization to precisely identify critical experts. Furthermore, an affine transformation-based expert approximation mechanism is designed to effectively compensate for accuracy degradation caused by pruning. Experimental results demonstrate that the proposed method successfully removes 25%–50% of experts across multiple large language models while consistently outperforming state-of-the-art algorithms on nine zero-shot benchmarks, achieving a favorable balance between model compression and reasoning capability.
Conventional pruning methods suffer from severe accuracy collapse at high sparsity levels, failing to meet stringent hardware constraints on model size. To address this, we propose a bidirectional pruning-regeneration framework that departs from traditional unidirectional pruning: it first applies aggressive structured pruning, then dynamically restores critical connections based on importance estimation and performance feedback. This iterative co-optimization of pruning and selective connection regeneration effectively mitigates accuracy degradation under extreme compression. Experiments demonstrate that our method achieves an average accuracy improvement of 4.2% over state-of-the-art approaches at equivalent sparsity levels. Notably, on ResNet-50, it attains 95% sparsity while retaining over 98% of the original accuracy—substantially outperforming existing pruning techniques. The proposed framework establishes a new paradigm for deploying highly accurate, ultra-sparse models on resource-constrained edge devices.
Existing structured pruning methods for large language models (LLMs) heavily rely on backpropagation, incurring substantial memory and computational overhead. To address this, we propose Bonsai—the first fully backpropagation-free, gradient-agnostic forward-pass pruning method for LLMs. Bonsai estimates module importance via forward perturbation analysis and performs module-level structured pruning without gradient computation. On a single NVIDIA A6000 GPU, Bonsai efficiently prunes the 8B-parameter LLaMA-3 model at 50% sparsity: memory consumption is reduced to one-half to one-third of conventional backward-based methods; pruning speed doubles; inference latency improves by 100%; and accuracy remains state-of-the-art. By eliminating dependence on gradient computation, Bonsai significantly broadens the feasibility of deploying compressed LLMs on resource-constrained hardware.
本文通过MoEXBench系统评估了组合压缩技术在Mixture-of-Experts模型中的应用,解决了部署难题,包括专家剪枝、权重量化和KV缓存压缩。
This work addresses the challenge of deploying Mixture-of-Experts (MoE) models, which suffer from high memory consumption and inference overhead. Existing compression methods apply coarse-grained pruning at the expert level, overlooking fine-grained redundancy within experts. To overcome this limitation, the paper proposes the first channel-level structured pruning framework for MoE models. It leverages attribution analysis to identify channels where information is concentrated and formulates the allocation of pruning ratios as a channel-score coverage maximization problem, which is efficiently solved to derive an optimal pruning strategy. Combined with 4-bit quantization, the method achieves nearly lossless accuracy under 50% or 25% structured pruning on DeepSeek and Qwen MoE models, respectively, and reduces memory usage by up to 5.27× on Qwen3-30B-A3B, significantly outperforming current state-of-the-art approaches.
This work addresses the challenge of efficiently compressing Mixture-of-Experts (MoE) models, which typically require loading all expert parameters. The authors propose a one-shot expert pruning method based on lightweight fine-tuning—such as router-specific LoRA or IA³—that induces changes in router weights. By measuring the ℓ² norm of these weight changes to assess expert sensitivity, they rank and prune the least sensitive experts. This study is the first to demonstrate that router sensitivity serves as an effective pruning signal, enabling near-linear accuracy degradation rather than catastrophic collapse under high compression ratios, with notable transferability across models. On Mixtral-8×7B, pruning 50% of experts yields a 28.76% score on MMLU-Pro, alongside 49% memory reduction and 37% lower latency; on Qwen1.5-MoE, it maintains a 49.7% average accuracy on mathematical tasks, substantially outperforming random or magnitude-based pruning.
This work addresses the computational, memory, and storage bottlenecks associated with deploying large-scale deep neural networks (DNNs) in resource-constrained environments. The authors propose a novel pruning method that integrates system-level engineering requirements with human-interpretable concepts—such as color and semantic categories—to identify critical neurons through analysis of their activation patterns, thereby guiding the generation of lightweight models. Notably, this approach is the first to incorporate interpretable concepts directly into the DNN pruning pipeline. Evaluated on VGG-19 using a dataset comprising 26,384 RGB images, the method yields pruned models that achieve substantial reductions in model size and computational overhead while maintaining high performance, demonstrating strong applicability across diverse real-world scenarios with stringent resource constraints.
研究通过动态专家剪枝方法在细粒度MoE架构中减少冗余专家选择,保留约2/3专家即可保持98.8%性能,提高推理效率。