Score
Designs and implements single-pass (one-shot) methods that jointly quantize and prune neural network components—removing redundant weights and reducing numeric precision of quantisers—in order to meet a target resource budget (memory, compute, or hardware resource counts) in a single operation while minimizing accuracy loss. Builds analyses and selection rules that align model size and performance to specified resource constraints by pruning both weights and quantisers once rather than via iterative retraining.
To address the trade-off among model size, inference latency, and accuracy degradation when deploying deep neural networks on edge devices, this paper proposes two co-designed pruning-quantization joint optimization frameworks. Methodologically, it tightly integrates feature-map similarity–based filter pruning with adaptive power-of-two (APoT) quantization, jointly optimizing pruning masks and low-bit (≤4-bit) quantization parameters during training. The key contribution lies in leveraging the complementarity of pruning and APoT: pruning eliminates structural redundancy, while APoT enhances quantized representation efficiency—thereby avoiding error accumulation inherent in sequential compression. Experiments on ResNet and VGG demonstrate that our approach achieves 5.2× model size reduction, 6.8× FLOPs reduction, and 4.1× inference speedup, with ≤0.3% Top-1 accuracy drop relative to full-precision baselines—substantially outperforming standalone pruning or quantization methods and exhibiting strong practicality for edge deployment.
The relative effectiveness of one-shot pruning versus iterative pruning across varying sparsity levels remains inadequately characterized, particularly under structured and unstructured pruning regimes. Method: We conduct a systematic benchmarking study across multiple pruning criteria, modalities, and sparsity ratios, providing the first rigorous formal definitions and comprehensive empirical comparison of both paradigms. We further propose a patience-based adaptive pruning criterion and a hybrid pruning strategy that synergistically integrates strengths of both approaches. Contribution/Results: Empirical results demonstrate that one-shot pruning achieves superior accuracy at low sparsity levels, whereas iterative pruning exhibits greater robustness under high sparsity. Our hybrid method consistently attains better accuracy–compression trade-offs across diverse tasks. This work establishes theoretical foundations and practical guidelines for selecting pruning strategies under real-world deployment constraints, advancing principled model compression.
In neural network quantized inference, conventional dot-product accumulation requires 32-bit accumulators to prevent overflow, resulting in high memory bandwidth and low energy efficiency. This work proposes PQS, the first framework to jointly integrate N:M floating-point pruning, ultra-low-bitweight/activation quantization (≤8 bits), and a “small-to-large” partial-sum ordering strategy—achieving algorithm–hardware co-design to eliminate accumulator overflow. By enabling extremely short accumulation chains, PQS reduces accumulator bit-width to just 12 bits—a 2.5× reduction—obviating the need for 32-bit accumulators. Evaluated across multiple image classification benchmarks, PQS maintains floating-point baseline accuracy while significantly improving memory bandwidth efficiency and system energy efficiency. The design exhibits strong hardware friendliness and practical deployability.
This work addresses the limitations of traditional high-granularity quantization (HGQ) in FPGA-based neural network compression, which relies on monotonic, irreversible layer-wise pruning that incurs substantial computational overhead and often fails to identify resource-constrained optimal subnetworks. To overcome this, the authors propose a resource-aware one-shot quantizer pruning method that directly maps the network into the target search space and integrates a bidirectional beta-scheduling fine-tuning strategy to efficiently explore the Pareto frontier between accuracy and hardware resource utilization. By eliminating progressive pruning, the approach reduces search cost by 20.58× compared to standard HGQ on jet substructure classification tasks while yielding a more competitive set of Pareto-optimal solutions and deployment configurations.
Neural network pruning under data-unavailable scenarios remains challenging, particularly in preserving model robustness without access to original training data. Method: This paper proposes a robustness-preserving pruning framework that operates without any training data. Departing from mainstream fine-tuning–dependent paradigms, it introduces robustness metrics—such as gradient sensitivity and adversarial response—as explicit supervision signals into the data-free pruning pipeline. The method integrates a progressive, conservative pruning strategy with random-optimization–driven channel-level sparsification to jointly optimize both accuracy and open-world robustness. Results: Extensive experiments across multiple CNN architectures demonstrate that our approach achieves an average 12.3% improvement in robust accuracy over state-of-the-art data-free pruning methods, while constraining clean accuracy degradation to within 1.5%. This significantly enhances the practicality and reliability of lightweight models deployed on resource-constrained devices.
This work addresses the discrepancy between conventional compression metrics—such as parameter count and FLOPs—and actual inference latency in CPU- and memory-constrained edge deployment scenarios, where such proxies often fail to reflect real-world performance. To bridge this gap, the authors propose a latency-driven, sequential compression pipeline that integrates unstructured pruning, INT8 quantization-aware training (QAT), and knowledge distillation (KD) within a unified training framework, jointly optimizing model accuracy, size, and inference speed. Experimental results demonstrate that this specific ordering significantly outperforms alternative combinations, achieving CPU inference latencies of 0.99–1.42 milliseconds on CIFAR-10/100 with ResNet-18, WRN-28-10, and VGG-16-BN models while maintaining high accuracy and compactness, thereby establishing a new paradigm for edge-oriented model compression under realistic latency constraints.
本文提出一种针对二值化神经网络的专用框架和全局加权剪枝算法,解决了现有剪枝策略不适用于二值化表示的问题,实现了更高的压缩率和精度平衡。
This work proposes a joint optimization pruning strategy that moves beyond the conventional paradigm of assessing weight importance solely based on magnitude. Recognizing that performance degradation caused by weight removal can be partially compensated through adjustments to adjacent biases, the method simultaneously prunes weights and computes optimal bias perturbations via automatic differentiation to minimize accuracy loss. By explicitly accounting for the interplay between weights and biases, this approach provides a more accurate measure of weight significance. Extensive experiments demonstrate that the proposed technique consistently outperforms state-of-the-art pruning methods across various models and tasks, achieving superior accuracy and robustness—particularly under high pruning ratios.
Quantized neural networks (QNNs) face challenges in achieving fully integer-only deployment due to their reliance on batch normalization (BN) layers, hindering efficient execution on resource-constrained edge devices. This paper proposes the first BN-free, end-to-end integer-only quantization framework. To replace BN’s functionality, we introduce inter-layer progressive knowledge distillation coupled with a dynamic compensation mechanism; further, we integrate parameter folding and custom integer-only operators to enable pure integer inference. Our method imposes no architectural or initialization constraints, significantly improving stability and accuracy under low-bit quantization. On ImageNet, our 4-bit quantized AlexNet achieves Top-1 accuracy within ±0.3% of the BN-based floating-point baseline—while eliminating all floating-point operations and runtime dependency on batch statistics. This represents the first practical solution for truly integer-only deployment of QNNs without BN.
研究通过一次性幅度剪枝和自适应提前退出方法减少神经网络计算量,并证明了这些方法的有效性及误差随计算差距的变化规律。