Bitwidth-Specific Logarithmic Arithmetic for Future Hardware-Accelerated Training

📅 2025-10-19
📈 Citations: 0
Influential: 0
📄 PDF

career value

173K/year
🤖 AI Summary
To address the high computational cost and low hardware efficiency of floating-point operations in deep learning training, this paper proposes a hardware-aware low-precision logarithmic fixed-point training method tailored for accelerators. The approach innovatively incorporates a bit-width–aware mechanism into logarithmic addition approximation, jointly optimizing piecewise linear approximation and simulated annealing to achieve Pareto-optimal trade-offs between accuracy and hardware overhead. Leveraging the logarithmic number system (LNS) and bit-accurate C++ simulation, end-to-end training is realized using 12-bit integer arithmetic. Experimental results on VGG-11 and VGG-16 demonstrate accuracy comparable to 32-bit floating-point training, while reducing multiply-accumulate (MAC) unit area by 32.5% and energy consumption by 53.5%. This work establishes a novel hardware–software co-design paradigm for low-precision deep learning training.

Technology Category

Application Category

📝 Abstract
While advancements in quantization have significantly reduced the computational costs of inference in deep learning, training still predominantly relies on complex floating-point arithmetic. Low-precision fixed-point training presents a compelling alternative. This work introduces a novel enhancement in low-precision logarithmic fixed-point training, geared towards future hardware accelerator designs. We propose incorporating bitwidth in the design of approximations to arithmetic operations. To this end, we introduce a new hardware-friendly, piece-wise linear approximation for logarithmic addition. Using simulated annealing, we optimize this approximation at different precision levels. A C++ bit-true simulation demonstrates training of VGG-11 and VGG-16 models on CIFAR-100 and TinyImageNet, respectively, using 12-bit integer arithmetic with minimal accuracy degradation compared to 32-bit floating-point training. Our hardware study reveals up to 32.5% reduction in area and 53.5% reduction in energy consumption for the proposed LNS multiply-accumulate units compared to that of linear fixed-point equivalents.
Problem

Research questions and friction points this paper is trying to address.

Enhancing low-precision logarithmic fixed-point training for hardware accelerators
Optimizing bitwidth-specific approximations for logarithmic arithmetic operations
Reducing area and energy in training with minimal accuracy loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bitwidth-specific logarithmic arithmetic for training
Hardware-friendly piece-wise linear approximation for addition
Optimized low-precision integer arithmetic with simulated annealing
🔎 Similar Papers
Hassan Hamad
Hassan Hamad
Ming Hsieh Department of Electrical and Computer Engineering, University of Southern California, Los Angeles, CA, USA
Y
Yuou Qiu
Ming Hsieh Department of Electrical and Computer Engineering, University of Southern California, Los Angeles, CA, USA
P
Peter A. Beerel
Ming Hsieh Department of Electrical and Computer Engineering, University of Southern California, Los Angeles, CA, USA
K
Keith M. Chugg
Ming Hsieh Department of Electrical and Computer Engineering, University of Southern California, Los Angeles, CA, USA