Pruning Binarized Neural Networks: A Dedicated Framework and Globally Weighted Algorithms

📅 2026-08-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种针对二值化神经网络的专用框架和全局加权剪枝算法,解决了现有剪枝策略不适用于二值化表示的问题,实现了更高的压缩率和精度平衡。
📝 Abstract
Extreme compression of deep neural networks, up to full binarization, dramatically reduces memory footprint and arithmetic complexity, facilitating deployment on constrained edge hardware with field-programmable gate arrays (FPGAs) and microcontrollers. Although combining binarization with pruning promises additional efficiency gains, existing pruning strategies are ill-suited to binarized representations and rarely translate into meaningful hardware savings. We introduce a PyTorch-based, research-oriented framework that incorporates freezing and pruning mechanisms for designing and optimizing binarized neural networks. The framework enables rapid and reproducible evaluation of state-of-the-art approaches and the fast prototyping of new ones. Leveraging this framework, we propose a novel pruning method that accounts for the relative importance of learned parameters across abstraction levels. Such a global weighting mechanism consistently achieves a superior trade-off between model accuracy and pruning rate, achieving a 70% pruning rate on VGG11 with constant accuracy, while state-of-the-art results reach only 41% in the binarized setting.
Problem

Research questions and friction points this paper is trying to address.

Binarized Neural Networks
Pruning
Hardware Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Binarized Neural Networks
Pruning
Global Weighting Mechanism
PyTorch Framework
🔎 Similar Papers
💼 Related Jobs
No related jobs found.