hardware-aware pruning

Designs, implements, and evaluates pruning and compression methods that remove redundant network weights, nodes, expert submodules, or circuit elements while explicitly optimizing hardware costs (area, power, latency, and cross-device communication) and preserving task performance. Builds cost and coverage metrics, task-aware or adaptive prior selection and transfer mechanisms, and importance- or coverage-based expert selection algorithms to decide which components to prune and to predict the resulting execution overhead and accuracy trade-offs.

hardware-awarepruning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.55
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$203K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

REAP the Experts: Why Pruning Prevails for One-Shot MoE compression

Oct 15, 2025
ML
Mike Lasby
🏛️ Cerebras Systems Inc. | University of Calgary

Sparse Mixture-of-Experts (MoE) models offer computational efficiency but suffer from functional subspace collapse during expert compression—especially in expert merging—where the router loses input-dependent control over expert selection, inducing irreversible errors. This work is the first to identify and formalize this deficiency. We propose Router-Weighted Expert Pruning (RWEP), a one-shot, fine-tuning-free compression method that jointly leverages gating weights and activation norms to dynamically assess expert importance. We theoretically prove that pruning preserves generative capability more faithfully than merging. Empirical evaluation on models ranging from 20B to 1T parameters shows that 50% expert pruning yields near-lossless performance: Qwen3-Coder-480B and Kimi-K2 achieve competitive code generation and tool-use accuracy, significantly outperforming both expert merging and alternative pruning baselines.

Merging causes functional subspace collapse in generative tasksREAP pruning outperforms merging for large-scale generative modelsSMoE models have excessive memory overhead requiring expert compression

This study addresses the substantial memory overhead of deploying Mixture-of-Experts (MoE) models and the limitations of existing pruning methods, including poor alignment, high computational cost, and neglect of routing redundancy. We propose an efficient MoE pruning framework that achieves structured pruning by introducing learnable router biases and diversity regularization to precisely identify critical experts. Furthermore, an affine transformation-based expert approximation mechanism is designed to effectively compensate for accuracy degradation caused by pruning. Experimental results demonstrate that the proposed method successfully removes 25%–50% of experts across multiple large language models while consistently outperforming state-of-the-art algorithms on nine zero-shot benchmarks, achieving a favorable balance between model compression and reasoning capability.

expert rankingmemory reductionMixture-of-Experts

Conventional pruning methods suffer from severe accuracy collapse at high sparsity levels, failing to meet stringent hardware constraints on model size. To address this, we propose a bidirectional pruning-regeneration framework that departs from traditional unidirectional pruning: it first applies aggressive structured pruning, then dynamically restores critical connections based on importance estimation and performance feedback. This iterative co-optimization of pruning and selective connection regeneration effectively mitigates accuracy degradation under extreme compression. Experiments demonstrate that our method achieves an average accuracy improvement of 4.2% over state-of-the-art approaches at equivalent sparsity levels. Notably, on ResNet-50, it attains 95% sparsity while retaining over 98% of the original accuracy—substantially outperforming existing pruning techniques. The proposed framework establishes a new paradigm for deploying highly accurate, ultra-sparse models on resource-constrained edge devices.

Addressing accuracy collapse beyond critical sparsity thresholdsEnabling extreme model compression for hardware constraintsOvercoming performance degradation in highly sparse neural networks

Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes

Feb 08, 2024
LD
L. Dery
🏛️ Carnegie Mellon University | Google Research

Existing structured pruning methods for large language models (LLMs) heavily rely on backpropagation, incurring substantial memory and computational overhead. To address this, we propose Bonsai—the first fully backpropagation-free, gradient-agnostic forward-pass pruning method for LLMs. Bonsai estimates module importance via forward perturbation analysis and performs module-level structured pruning without gradient computation. On a single NVIDIA A6000 GPU, Bonsai efficiently prunes the 8B-parameter LLaMA-3 model at 50% sparsity: memory consumption is reduced to one-half to one-third of conventional backward-based methods; pruning speed doubles; inference latency improves by 100%; and accuracy remains state-of-the-art. By eliminating dependence on gradient computation, Bonsai significantly broadens the feasibility of deploying compressed LLMs on resource-constrained hardware.

Achieve high sparsity pruning on large models with limited GPU resourcesDevelop gradient-free pruning for LLMs to reduce memory and compute costsEnable efficient model compression on diverse hardware without backpropagation

Latest Papers

What's happening recently
View more

This work addresses the challenge of deploying Mixture-of-Experts (MoE) models, which suffer from high memory consumption and inference overhead. Existing compression methods apply coarse-grained pruning at the expert level, overlooking fine-grained redundancy within experts. To overcome this limitation, the paper proposes the first channel-level structured pruning framework for MoE models. It leverages attribution analysis to identify channels where information is concentrated and formulates the allocation of pruning ratios as a channel-score coverage maximization problem, which is efficiently solved to derive an optimal pruning strategy. Combined with 4-bit quantization, the method achieves nearly lossless accuracy under 50% or 25% structured pruning on DeepSeek and Qwen MoE models, respectively, and reduces memory usage by up to 5.27× on Qwen3-30B-A3B, significantly outperforming current state-of-the-art approaches.

memory footprintMixture-of-Expertsmodel compression

This work addresses the challenge of efficiently compressing Mixture-of-Experts (MoE) models, which typically require loading all expert parameters. The authors propose a one-shot expert pruning method based on lightweight fine-tuning—such as router-specific LoRA or IA³—that induces changes in router weights. By measuring the ℓ² norm of these weight changes to assess expert sensitivity, they rank and prune the least sensitive experts. This study is the first to demonstrate that router sensitivity serves as an effective pruning signal, enabling near-linear accuracy degradation rather than catastrophic collapse under high compression ratios, with notable transferability across models. On Mixtral-8×7B, pruning 50% of experts yields a 28.76% score on MMLU-Pro, alongside 49% memory reduction and 37% lower latency; on Qwen1.5-MoE, it maintains a 49.7% average accuracy on mathematical tasks, substantially outperforming random or magnitude-based pruning.

expert pruninglightweight fine-tuningMixture-of-Experts

This work addresses the computational, memory, and storage bottlenecks associated with deploying large-scale deep neural networks (DNNs) in resource-constrained environments. The authors propose a novel pruning method that integrates system-level engineering requirements with human-interpretable concepts—such as color and semantic categories—to identify critical neurons through analysis of their activation patterns, thereby guiding the generation of lightweight models. Notably, this approach is the first to incorporate interpretable concepts directly into the DNN pruning pipeline. Evaluated on VGG-19 using a dataset comprising 26,384 RGB images, the method yields pruned models that achieve substantial reductions in model size and computational overhead while maintaining high performance, demonstrating strong applicability across diverse real-world scenarios with stringent resource constraints.

Deep Neural Networksmodel pruningresource-constrained systems

Hot Scholars

TY

Tao Yuan

University of California, Los Angeles
Computer VisionArtificial Intelligence
CY

Changdi Yang

PhD candidate, Northeastern University, Snap Inc.
Efficient Deep Learning
GL

Guiying Li

Pengcheng Laboratory
DNN compressionCloud NativeEdge ComputingStock Market
HY

Haoran Yang

Central South University
Graph Neural NetworksData MiningRecommendation Systems