perform model pruning

Designs, implements, and evaluates methods that remove parameters or structure from machine learning models to reduce memory, compute, or latency while preserving performance; this includes magnitude-based and algorithmic parameter pruning, development of pruning strategies and schedules, and techniques for reintroducing capacity. It also covers formulating and searching pruning search spaces and sparsity patterns, integrating pruning with other compressions (e.g., quantization), and analyzing the effects of different pruning algorithms.

performmodelpruning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$225K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a joint optimization pruning strategy that moves beyond the conventional paradigm of assessing weight importance solely based on magnitude. Recognizing that performance degradation caused by weight removal can be partially compensated through adjustments to adjacent biases, the method simultaneously prunes weights and computes optimal bias perturbations via automatic differentiation to minimize accuracy loss. By explicitly accounting for the interplay between weights and biases, this approach provides a more accurate measure of weight significance. Extensive experiments demonstrate that the proposed technique consistently outperforms state-of-the-art pruning methods across various models and tasks, achieving superior accuracy and robustness—particularly under high pruning ratios.

bias compensationneural networkspruning

Conventional pruning methods suffer from severe accuracy collapse at high sparsity levels, failing to meet stringent hardware constraints on model size. To address this, we propose a bidirectional pruning-regeneration framework that departs from traditional unidirectional pruning: it first applies aggressive structured pruning, then dynamically restores critical connections based on importance estimation and performance feedback. This iterative co-optimization of pruning and selective connection regeneration effectively mitigates accuracy degradation under extreme compression. Experiments demonstrate that our method achieves an average accuracy improvement of 4.2% over state-of-the-art approaches at equivalent sparsity levels. Notably, on ResNet-50, it attains 95% sparsity while retaining over 98% of the original accuracy—substantially outperforming existing pruning techniques. The proposed framework establishes a new paradigm for deploying highly accurate, ultra-sparse models on resource-constrained edge devices.

Addressing accuracy collapse beyond critical sparsity thresholdsEnabling extreme model compression for hardware constraintsOvercoming performance degradation in highly sparse neural networks

One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression

Aug 19, 2025
MJ
Mikołaj Janusz
🏛️ Jagiellonian University | ETH Zürich | Wrocław University of Science and Technology

The relative effectiveness of one-shot pruning versus iterative pruning across varying sparsity levels remains inadequately characterized, particularly under structured and unstructured pruning regimes. Method: We conduct a systematic benchmarking study across multiple pruning criteria, modalities, and sparsity ratios, providing the first rigorous formal definitions and comprehensive empirical comparison of both paradigms. We further propose a patience-based adaptive pruning criterion and a hybrid pruning strategy that synergistically integrates strengths of both approaches. Contribution/Results: Empirical results demonstrate that one-shot pruning achieves superior accuracy at low sparsity levels, whereas iterative pruning exhibits greater robustness under high sparsity. Our hybrid method consistently attains better accuracy–compression trade-offs across diverse tasks. This work establishes theoretical foundations and practical guidelines for selecting pruning strategies under real-world deployment constraints, advancing principled model compression.

Evaluates pruning effectiveness across structured and unstructured compression settingsIdentifies optimal pruning strategies based on specific compression ratiosSystematically compares one-shot and iterative neural network pruning methods

PERP: Rethinking the Prune-Retrain Paradigm in the Era of LLMs

Dec 23, 2023
MZ
Max Zimmer
🏛️ Zuse Institute Berlin | Technische Universität Berlin

To address the infeasibility of full retraining after pruning large language models (LLMs) due to GPU memory and computational constraints, this paper proposes a retraining-free, extremely sparse fine-tuning paradigm: updating only 0.01%–0.05% of the most expressive parameters suffices to restore—or even surpass—full-retraining performance. Key contributions include: (1) the first empirical demonstration that ultra-low-parameter updates can effectively substitute full retraining; (2) two mergeable, sparsity-preserving LoRA variants; and (3) an inter-layer weight reconstruction mechanism for efficient sparse model enhancement. On GPT architectures, pruning followed by fine-tuning completes in minutes on a single GPU for a 30B model. Across sparsity levels, our method matches or exceeds full retraining accuracy and significantly outperforms baseline approaches—including Wanda and SparseGPT—while maintaining structural sparsity and computational efficiency.

Enhances performance with minimal parameter updatesOptimizes neural network pruning efficiencyReduces retraining computational demands

Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes

Feb 08, 2024
LD
L. Dery
🏛️ Carnegie Mellon University | Google Research

Existing structured pruning methods for large language models (LLMs) heavily rely on backpropagation, incurring substantial memory and computational overhead. To address this, we propose Bonsai—the first fully backpropagation-free, gradient-agnostic forward-pass pruning method for LLMs. Bonsai estimates module importance via forward perturbation analysis and performs module-level structured pruning without gradient computation. On a single NVIDIA A6000 GPU, Bonsai efficiently prunes the 8B-parameter LLaMA-3 model at 50% sparsity: memory consumption is reduced to one-half to one-third of conventional backward-based methods; pruning speed doubles; inference latency improves by 100%; and accuracy remains state-of-the-art. By eliminating dependence on gradient computation, Bonsai significantly broadens the feasibility of deploying compressed LLMs on resource-constrained hardware.

Achieve high sparsity pruning on large models with limited GPU resourcesDevelop gradient-free pruning for LLMs to reduce memory and compute costsEnable efficient model compression on diverse hardware without backpropagation

Latest Papers

What's happening recently
View more

Existing data pruning methods suffer significant performance degradation under high label noise and struggle to effectively retain informative samples. This work systematically investigates the behavior of pruning strategies in both noisy and noise-free settings, and for the first time explicitly identifies data redundancy, problematic samples, and inter-sample dependencies as three universal factors governing pruning efficacy. Through empirical analysis of two dominant pruning paradigms across standard classification benchmarks and mainstream neural architectures, the study demonstrates the consistent influence of these factors under diverse data distributions and training protocols. The findings not only expose fundamental limitations of current approaches but also offer a new perspective toward designing robust pruning methods.

data pruninglabel noiseproblematic samples

Pruning and Quantization Impact on Graph Neural Networks

Oct 24, 2025
KK
Khatoon Khedri
🏛️ Boston University | Chulalongkorn University | Duy Tan University | Galgotias University

Graph Neural Networks (GNNs) achieve high accuracy but incur substantial computational and memory overhead, hindering deployment on resource-constrained devices. Method: This work systematically investigates the effectiveness of pruning and quantization for GNN compression across node classification (Cora), graph classification (Proteins), and link prediction (BBBP). We integrate unstructured fine-grained pruning, global pruning, and structured pruning with fixed-point quantization, dynamic quantization, and mixed-precision quantization, complemented by lightweight fine-tuning. Contribution/Results: Unstructured fine-grained and global pruning maintain or even improve accuracy while reducing model size by 50%. Quantization methods exhibit strong dataset dependence, revealing clear trade-offs between inference latency and model footprint. Crucially, we empirically identify a “sparsity–accuracy positive correlation” phenomenon in GNN pruning—contrary to conventional wisdom—providing the first empirical evidence and practical recipes for efficient GNN deployment.

Analyzing model size reduction while preserving graph learning performanceEvaluating compression techniques for maintaining GNN accuracy across tasksInvestigating pruning and quantization effects on GNN computational efficiency

Existing research indicates that structured pruning substantially degrades the test-time scaling (TTS) performance of large language models, yet it remains unclear how to compress models while preserving or even enhancing TTS capabilities. This work systematically investigates the impact of unstructured pruning on TTS, conducting experiments on the S1.1-7B and Qwen3-8B models with various inter-layer sparsity allocation strategies. The results demonstrate that unstructured pruning not only effectively mitigates performance degradation but also surpasses the original unpruned models across multiple reasoning benchmarks, significantly outperforming structured pruning. These findings challenge the prevailing assumption that pruning inevitably harms TTS and offer a promising new direction for efficient model compression that supports high-performance inference.

inference efficiencyLLM pruningreasoning performance

Dataset Pruning in RecSys and ML: Best Practice or Mal-Practice?

Oct 16, 2025
LW
Leonie Winter
🏛️ University of Siegen

User interaction frequency pruning—commonly applied to filter out low-activity users—introduces systematic biases in dataset characteristics and algorithm evaluation in recommender systems. Method: We conduct a comprehensive empirical study across five public datasets, each augmented with multiple levels of pruning. We train and evaluate 11 state-of-the-art recommendation algorithms under both standard offline evaluation and cross-pruning generalization settings. Contribution/Results: We quantitatively demonstrate that conventional pruning severely reduces user coverage, inflating model performance on pruned test sets while consistently degrading performance on unpruned, full-distribution test sets. The magnitude of evaluation distortion increases monotonically with pruning intensity. This work is the first to formally quantify this “evaluation distortion” phenomenon induced by data pruning. We advocate for cautious adoption of pruning heuristics and emphasize robust evaluation under the complete, unaltered data distribution. Our findings provide critical methodological guidance for establishing reliable experimental benchmarks in recommender system research.

Analyzing performance decline when models trained on pruned data face unpruned testsAssessing if pruned datasets artificially inflate recommender system performance metricsEvaluating how data pruning affects dataset characteristics and algorithm performance

This work proposes a novel pruning paradigm grounded in selection dynamics, addressing the limitations of conventional neural network pruning methods that rely on explicit intervention and centralized decision-making—approaches ill-suited to the decentralized, stochastic, and path-dependent nature of gradient-based training. By conceptualizing neurons as an evolving population subject to selection pressure, the method defines neuronal fitness through local learning signals, allowing low-fitness components to naturally vanish during training without predefined pruning schedules. Sparsity thus emerges intrinsically as an evolutionary outcome of training rather than being imposed as an external constraint. Evaluated on MNIST with a 768-neuron prunable MLP, the approach achieves 95.5% accuracy at 35% pruning and maintains 88.3–88.6% accuracy at 50% pruning, closely approaching the 98% performance of the original dense model.

evolutionary dynamicsneural networkspruning

Hot Scholars

MZ

Max Zimmer

Zuse Institute Berlin
Deep LearningOptimizationMathematics
CR

Christophe Roux

TU Berlin, Zuse Institute Berlin
OptimizationMachine Learning
XL

Xunkai Li

School of Computer Science and Technology, Beijing Institution of Technology
Data-centric AIGraph MLAI4Science
DA

Dan Alistarh

Professor at IST Austria
Machine LearningAlgorithmsDistributed Computing
QD

Qiangqiang Dai

Beijing Institute of Technology
Graph Data Management