one-shot quantization and pruning

Designs and implements single-pass (one-shot) methods that jointly quantize and prune neural network components—removing redundant weights and reducing numeric precision of quantisers—in order to meet a target resource budget (memory, compute, or hardware resource counts) in a single operation while minimizing accuracy loss. Builds analyses and selection rules that align model size and performance to specified resource constraints by pruning both weights and quantisers once rather than via iterative retraining.

one-shotquantizationandpruning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address the trade-off among model size, inference latency, and accuracy degradation when deploying deep neural networks on edge devices, this paper proposes two co-designed pruning-quantization joint optimization frameworks. Methodologically, it tightly integrates feature-map similarity–based filter pruning with adaptive power-of-two (APoT) quantization, jointly optimizing pruning masks and low-bit (≤4-bit) quantization parameters during training. The key contribution lies in leveraging the complementarity of pruning and APoT: pruning eliminates structural redundancy, while APoT enhances quantized representation efficiency—thereby avoiding error accumulation inherent in sequential compression. Experiments on ResNet and VGG demonstrate that our approach achieves 5.2× model size reduction, 6.8× FLOPs reduction, and 4.1× inference speedup, with ≤0.3% Top-1 accuracy drop relative to full-precision baselines—substantially outperforming standalone pruning or quantization methods and exhibiting strong practicality for edge deployment.

Addressing computational and memory demands on resource-constrained devicesCombining pruning and quantization for efficient DNN compressionPreserving model accuracy while achieving higher compression efficiency

One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression

Aug 19, 2025
MJ
Mikołaj Janusz
🏛️ Jagiellonian University | ETH Zürich | Wrocław University of Science and Technology

The relative effectiveness of one-shot pruning versus iterative pruning across varying sparsity levels remains inadequately characterized, particularly under structured and unstructured pruning regimes. Method: We conduct a systematic benchmarking study across multiple pruning criteria, modalities, and sparsity ratios, providing the first rigorous formal definitions and comprehensive empirical comparison of both paradigms. We further propose a patience-based adaptive pruning criterion and a hybrid pruning strategy that synergistically integrates strengths of both approaches. Contribution/Results: Empirical results demonstrate that one-shot pruning achieves superior accuracy at low sparsity levels, whereas iterative pruning exhibits greater robustness under high sparsity. Our hybrid method consistently attains better accuracy–compression trade-offs across diverse tasks. This work establishes theoretical foundations and practical guidelines for selecting pruning strategies under real-world deployment constraints, advancing principled model compression.

Evaluates pruning effectiveness across structured and unstructured compression settingsIdentifies optimal pruning strategies based on specific compression ratiosSystematically compares one-shot and iterative neural network pruning methods

In neural network quantized inference, conventional dot-product accumulation requires 32-bit accumulators to prevent overflow, resulting in high memory bandwidth and low energy efficiency. This work proposes PQS, the first framework to jointly integrate N:M floating-point pruning, ultra-low-bitweight/activation quantization (≤8 bits), and a “small-to-large” partial-sum ordering strategy—achieving algorithm–hardware co-design to eliminate accumulator overflow. By enabling extremely short accumulation chains, PQS reduces accumulator bit-width to just 12 bits—a 2.5× reduction—obviating the need for 32-bit accumulators. Evaluated across multiple image classification benchmarks, PQS maintains floating-point baseline accuracy while significantly improving memory bandwidth efficiency and system energy efficiency. The design exhibits strong hardware friendliness and practical deployability.

Maintaining accuracy while compressing neural network modelsPreventing overflow in low-bitwidth dot product accumulationReducing accumulator bitwidth in neural network computations

This work addresses the limitations of traditional high-granularity quantization (HGQ) in FPGA-based neural network compression, which relies on monotonic, irreversible layer-wise pruning that incurs substantial computational overhead and often fails to identify resource-constrained optimal subnetworks. To overcome this, the authors propose a resource-aware one-shot quantizer pruning method that directly maps the network into the target search space and integrates a bidirectional beta-scheduling fine-tuning strategy to efficiently explore the Pareto frontier between accuracy and hardware resource utilization. By eliminating progressive pruning, the approach reduces search cost by 20.58× compared to standard HGQ on jet substructure classification tasks while yielding a more competitive set of Pareto-optimal solutions and deployment configurations.

FPGAsPareto frontierpruning

Supervised Robustness-preserving Data-free Neural Network Pruning

Apr 02, 2022
MH
M. H. Meng
🏛️ National University of Singapore | Institute for Infocomm Research | A*STAR | The University of Queensland

Neural network pruning under data-unavailable scenarios remains challenging, particularly in preserving model robustness without access to original training data. Method: This paper proposes a robustness-preserving pruning framework that operates without any training data. Departing from mainstream fine-tuning–dependent paradigms, it introduces robustness metrics—such as gradient sensitivity and adversarial response—as explicit supervision signals into the data-free pruning pipeline. The method integrates a progressive, conservative pruning strategy with random-optimization–driven channel-level sparsification to jointly optimize both accuracy and open-world robustness. Results: Extensive experiments across multiple CNN architectures demonstrate that our approach achieves an average 12.3% improvement in robust accuracy over state-of-the-art data-free pruning methods, while constraining clean accuracy degradation to within 1.5%. This significantly enhances the practicality and reliability of lightweight models deployed on resource-constrained devices.

Develops data-free pruning for resource-constrained devicesEnsures pruned models maintain accuracy and robustnessReplaces aggressive one-shot pruning with progressive optimization

Latest Papers

What's happening recently
View more

This work addresses the discrepancy between conventional compression metrics—such as parameter count and FLOPs—and actual inference latency in CPU- and memory-constrained edge deployment scenarios, where such proxies often fail to reflect real-world performance. To bridge this gap, the authors propose a latency-driven, sequential compression pipeline that integrates unstructured pruning, INT8 quantization-aware training (QAT), and knowledge distillation (KD) within a unified training framework, jointly optimizing model accuracy, size, and inference speed. Experimental results demonstrate that this specific ordering significantly outperforms alternative combinations, achieving CPU inference latencies of 0.99–1.42 milliseconds on CIFAR-10/100 with ResNet-18, WRN-28-10, and VGG-16-BN models while maintaining high accuracy and compactness, thereby establishing a new paradigm for edge-oriented model compression under realistic latency constraints.

inference latencyknowledge distillationneural network compression

This work proposes a joint optimization pruning strategy that moves beyond the conventional paradigm of assessing weight importance solely based on magnitude. Recognizing that performance degradation caused by weight removal can be partially compensated through adjustments to adjacent biases, the method simultaneously prunes weights and computes optimal bias perturbations via automatic differentiation to minimize accuracy loss. By explicitly accounting for the interplay between weights and biases, this approach provides a more accurate measure of weight significance. Extensive experiments demonstrate that the proposed technique consistently outperforms state-of-the-art pruning methods across various models and tasks, achieving superior accuracy and robustness—particularly under high pruning ratios.

bias compensationneural networkspruning

Batch Normalization-Free Fully Integer Quantized Neural Networks via Progressive Tandem Learning

Dec 18, 2025
PS
Pengfei Sun
🏛️ Ghent University | Institute for Infocomm Research | Agency for Science, Technology and Research

Quantized neural networks (QNNs) face challenges in achieving fully integer-only deployment due to their reliance on batch normalization (BN) layers, hindering efficient execution on resource-constrained edge devices. This paper proposes the first BN-free, end-to-end integer-only quantization framework. To replace BN’s functionality, we introduce inter-layer progressive knowledge distillation coupled with a dynamic compensation mechanism; further, we integrate parameter folding and custom integer-only operators to enable pure integer inference. Our method imposes no architectural or initialization constraints, significantly improving stability and accuracy under low-bit quantization. On ImageNet, our 4-bit quantized AlexNet achieves Top-1 accuracy within ±0.3% of the BN-based floating-point baseline—while eliminating all floating-point operations and runtime dependency on batch statistics. This represents the first practical solution for truly integer-only deployment of QNNs without BN.

Eliminates batch normalization for integer-only neural networksEnables fully integer inference via progressive layer-wise distillationMaintains accuracy in aggressive quantization for edge devices

Hot Scholars

JZ

Jingren Zhou

Alibaba Group, Microsoft
Cloud ComputingLarge Scale Distributed SystemsMachine LearningQuery Processing
CC

Cheng Cui

BUAA
deep learningnetwork designOCRmllm
JW

Jun-Wei Hsieh

National Yang Ming Chiao Tung University
computer visionAIimage processing
JQ

Jinwei Qi

Tongyi Lab, Alibaba Group
Artificial Intelligencedeep learningmultimedia
KW

Kezhi Wang

Professor, Royal Society Industry Fellow, Brunel University London
Wireless CommunicationEdge ComputingMachine Learning