🤖 AI Summary
This work addresses the challenges of high inference latency, substantial memory consumption, and excessive power usage in convolutional neural networks (CNNs) due to their large scale, as well as the lack of a unified framework and comparability among existing pruning methods. To this end, the authors propose Bonsai, a novel pruning framework that establishes a general-purpose platform called Combine, introduces a standardized language for describing pruning criteria, and incorporates several new filter-level pruning strategies. Through an iterative structured pruning approach, Bonsai removes up to 79% of filters in VGG-style models, reduces computational cost by 68%, and maintains or even improves model accuracy. The study systematically elucidates the performance disparities arising from different pruning criteria.
📝 Abstract
As the need for more accurate and powerful Convolutional Neural Networks (CNNs) increases, so too does the size, execution time, memory footprint, and power consumption. To overcome this, solutions such as pruning have been proposed with their own metrics and methodologies, or criteria, for how weights should be removed. These solutions do not share a common implementation and are difficult to implement and compare. In this work, we introduce Combine, a criterion- based pruning solution and demonstrate that it is fast and effective framework for iterative pruning, demonstrate that criterion have differing effects on different models, create a standard language for comparing criterion functions, and propose a few novel criterion functions. We show the capacity of these criterion functions and the framework on VGG inspired models, pruning up to 79\% of filters while retaining or improving accuracy, and reducing the computations needed by the network by up to 68\%.