structural pruning

Designs and implements methods that remove or disable whole model components—such as channels, filters, layers, and attention heads—by imposing structured sparsity or inserting binary/per-head gates so the remaining model structure is preserved. Builds end-to-end optimization and evaluation pipelines that trade off compute/memory (FLOPs) against accuracy, using techniques like distillation and noise-weighted objectives, and analyzes accuracy degradation and compression behavior under high sparsity.

structuralpruning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.77
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$246K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Sparse Computations in Deep Learning Inference

Dec 02, 2025
IT
Ioanna Tasou
🏛️ National Technical University of Athens | Max Planck Institute for Software Systems | University of Glasgow

The computational and energy overhead of deep learning inference is increasingly prohibitive, yet sparsity—a key optimization avenue—remains underutilized in production systems. Method: Targeting performance engineers, this work systematically surveys structured and unstructured sparsity exploitable in DNN inference and proposes an end-to-end engineering methodology—from sparse model representation to efficient sparse kernels (SpMM/SDDMM). We implement and benchmark multiple sparse computation schemes on CPU and GPU platforms, integrating support for mainstream frameworks, toolchains, and datasets. Contribution/Results: We present the first production-grade sparse inference reference framework encompassing hardware adaptation, kernel optimization, and deployment validation. Experiments demonstrate 2–5× inference speedup and substantial energy efficiency gains across representative models, establishing a reproducible, scalable practical paradigm for industrial deployment of sparse deep learning.

Bridging the gap between sparsity potential and production AI systemsOptimizing sparse computations for efficient deep learning inferenceProviding knowledge for performance engineers on sparse DNN implementation

Effective Interplay between Sparsity and Quantization: From Theory to Practice

May 31, 2024
SB
Simla Burcu Harma
🏛️ EPFL | MangoBoost Inc. | Google | Korea University | Google DeepMind

To address efficient deployment of large language and vision models on resource-constrained devices, this work systematically investigates the non-orthogonality between sparsification and quantization—two key model compression techniques—and characterizes their joint accuracy degradation mechanism. We provide the first rigorous mathematical proof of their non-orthogonality, revealing that error accumulation is intrinsic and critically dependent on operational ordering: quantizing before sparsifying severely degrades accuracy, whereas sparsifying before quantizing mitigates error propagation. Grounded in theoretical analysis and validated across diverse models (OPT, LLaMA-125M–8B, ViT, ResNet) and hardware platforms, we establish “sparsify-then-quantize” as the optimal practice. This principle achieves high compression ratios while markedly improving the accuracy–efficiency trade-off, offering both theoretically grounded insights and a practical, empirically verified paradigm for edge deployment of foundation models.

Model CompressionResource-constrained DevicesSparse Quantization

MSQ: Memory-Efficient Bit Sparsification Quantization

Jul 29, 2025
SH
Seokho Han
🏛️ Sungkyunkwan University | University of Arizona

To address efficiency bottlenecks in deploying deep neural networks (DNNs) on mobile and edge devices, this paper proposes an end-to-end trainable mixed-precision quantization method. The approach introduces a differentiable round-clamp quantizer that jointly leverages Hessian information to enforce bit-level sparsity regularization. Crucially, it avoids explicit bit-width parameter separation and instead performs differentiable pruning directly on weight低位 bits, unifying precision allocation and sparse structure optimization. This design circumvents the high memory overhead and training complexity inherent in conventional bit-level sparse methods. Experiments demonstrate that the method achieves state-of-the-art accuracy and compression ratios while reducing trainable parameters by up to 8.00× and accelerating training time by 86%, significantly enhancing edge-device adaptability and deployment efficiency.

Enhancing precision reduction without bit-level splittingOptimizing mixed-precision quantization for DNN efficiencyReducing training complexity and GPU memory usage

Conventional pruning methods suffer from severe accuracy collapse at high sparsity levels, failing to meet stringent hardware constraints on model size. To address this, we propose a bidirectional pruning-regeneration framework that departs from traditional unidirectional pruning: it first applies aggressive structured pruning, then dynamically restores critical connections based on importance estimation and performance feedback. This iterative co-optimization of pruning and selective connection regeneration effectively mitigates accuracy degradation under extreme compression. Experiments demonstrate that our method achieves an average accuracy improvement of 4.2% over state-of-the-art approaches at equivalent sparsity levels. Notably, on ResNet-50, it attains 95% sparsity while retaining over 98% of the original accuracy—substantially outperforming existing pruning techniques. The proposed framework establishes a new paradigm for deploying highly accurate, ultra-sparse models on resource-constrained edge devices.

Addressing accuracy collapse beyond critical sparsity thresholdsEnabling extreme model compression for hardware constraintsOvercoming performance degradation in highly sparse neural networks

This study addresses the high inference computational cost and limited sparse activation of feedforward networks by proposing model casting and low-parameter gating methods. By employing highly sparse gating matrices to bypass redundant computations and introducing an innovative low-FLOPs parameterization design, the proposed approach surpasses the 3× acceleration ceiling of standard gating techniques. Furthermore, it achieves efficient inference by integrating mid-training strategies, sparse activation functions, and dedicated GPU kernel optimizations. Experimental results demonstrate that the method maintains comparable generation quality while attaining a theoretical speedup of 3.2× and an empirical GPU speedup of 3.31×, significantly outperforming existing approaches.

Feed-Forward NetworkFLOPs reductioninference acceleration

Latest Papers

What's happening recently
View more

This work addresses a critical mismatch between existing diffusion model acceleration techniques—which rely on element-wise activation sparsity—and the column-granularity data processing inherent in modern hardware, leading to overestimated practical sparsity benefits. The study presents the first systematic characterization of output sparsity across seven diffusion models at the column level, identifying three distinct activation distribution patterns and uncovering the interplay between model architecture and memory layout optimization. Through hardware-aware column-level sparsity profiling, cycle-accurate GDDR6 simulation, multi-threshold accuracy evaluation, and cross-modal comparison, the authors demonstrate that memory stalls account for 84–89% of total execution cycles and that element-wise sparsity poorly predicts actual hardware gains. Their approach achieves up to 30.6% reduction in execution cycles on UNet+Transformer models, with a maximum latency decrease (MLD) of 50.8%.

activation sparsitycolumn-level sparsitydiffusion models

This work addresses the challenge of unifying diverse structured sparsity patterns for efficient model compression and acceleration. The authors propose S³, an algebraic framework that formally integrates three core components—View (tensor reshaping), Block (atomic pruning units), and Scope (sparsity decision range)—to express a wide spectrum of sparsity patterns, ranging from fine-grained N:M sparsity to coarse-grained channel pruning, within a single formalism. Notably, S³ enables cross-tensor collaborative sparsification. Building upon this framework, the authors incorporate Optimal Brain Damage and Surgeon algorithms to develop structured variants of OBS/OBD. These methods significantly outperform current state-of-the-art second-order heuristic approaches in terms of output reconstruction accuracy.

model compressionsparse patternssparsity specification

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

This work addresses the irregular memory access induced by weight sparsity in event-driven SIMD/SIMT neuromorphic computing, which incurs a “sparsity tax” manifested in increased control complexity, metadata overhead, and reduced memory efficiency. For the first time, this study quantifies this sparsity tax and proposes three closely related accelerator architectures: a baseline SIMD, a bitmap-gated Sparse-SIMD, and an SIMT design leveraging run-length encoded (RLE) sparse weights. Evaluated via RTL-to-gate implementation in GF22FDX+ technology combined with an activity-driven energy model, results show that SRAM-dominated area limits total area variation; SIMT achieves substantial gains in energy efficiency and performance at high sparsity levels, albeit constrained by metadata bandwidth and workload imbalance; Sparse-SIMD yields only modest improvements due to bitmap and dense-storage overheads.

event-driven neuromorphicSIMDSIMT

Hot Scholars

PM

Pavlo Molchanov

NVIDIA Research
AIMachine LearningEfficient Deep LearningSemi-supervised learning
SH

Shwai He

University of Maryland, College Park
Deep LearningMechine LearningNatural Language Processing.
HY

Huanrui Yang

Assistant Professor, ECE, University of Arizona
Efficient deep learningTrustworthy deep learning
KK

Kenji Kawaguchi

Presidential Young Professor, National University of Singapore
LLMsLarge language modelDeep learningAI