Hardware-Efficient CNNs: Interleaved Approximate FP32 Multipliers for Kernel Computation

๐Ÿ“… 2025-10-08
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Conventional FP32 multipliers incur substantial hardware overhead and poor energy efficiency, hindering high-performance CNN inference. Method: This paper proposes a co-optimization framework for deploying error-bounded approximate FP32 multipliers *within* convolution kernels, targeting CNN inference. Leveraging the IEEE 754 standard, it employs approximate compressors for significand multiplication and integrates NSGA-IIโ€”a multi-objective genetic algorithmโ€”to jointly optimize multiplier type selection, placement, and composition order, thereby balancing accuracy and hardware efficiency. Contribution/Results: Evaluated across multiple CNN models, the approach maintains >99% of the original accuracy while reducing multiplier area and power consumption by 32.7% and 28.4% on average, respectively. This yields significant improvements in inference energy efficiency and throughput, providing a deployable, architecture-level solution for high-precision approximate computing.

Technology Category

Machine Learning: Hardware-aware MLComputer Vision: Learning & Optimization for CVSearch and Optimization: Non-convex Optimization

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphs
๐Ÿ“ Abstract
Single-precision floating point (FP32) data format, defined by the IEEE 754 standard, is widely employed in scientific computing, signal processing, and deep learning training, where precision is critical. However, FP32 multiplication is computationally expensive and requires complex hardware, especially for precisely handling mantissa multiplication. In practical applications like neural network inference, perfect accuracy is not always necessary, minor multiplication errors often have little impact on final accuracy. This enables trading precision for gains in area, power, and speed. This work focuses on CNN inference using approximate FP32 multipliers, where the mantissa multiplication is approximated by employing error-variant approximate compressors, that significantly reduce hardware cost. Furthermore, this work optimizes CNN performance by employing differently approximated FP32 multipliers and studying their impact when interleaved within the kernels across the convolutional layers. The placement and ordering of these approximate multipliers within each kernel are carefully optimized using the Non-dominated Sorting Genetic Algorithm-II, balancing the trade-off between accuracy and hardware efficiency.
Problem

Research questions and friction points this paper is trying to address.

Designing approximate FP32 multipliers to reduce hardware costs
Optimizing CNN performance with interleaved approximate multipliers
Balancing accuracy and hardware efficiency using genetic algorithms
Innovation

Methods, ideas, or system contributions that make the work stand out.

Approximate FP32 multipliers reduce hardware cost
Interleaved multipliers optimize CNN performance trade-offs
Genetic algorithm optimizes multiplier placement for efficiency
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
B
Bindu G Gowda
IIIT-Bangalore, India
Yogesh Goyal
Yogesh Goyal
IIIT-Bangalore, India
Yash Gupta
Yash Gupta
IIIT-Bangalore, India
M
Madhav Rao
IIIT-Bangalore, India