A flexible FPGA accelerator for convolutional neural networks

📅 2019-12-16
🏛️ arXiv.org
📈 Citations: 5
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the memory bandwidth bottleneck and hardware programming complexity in FPGA-accelerated CNN inference, this paper proposes a scalable software-hardware co-design acceleration architecture. Methodologically: (1) it introduces a reconfiguration-free 1D processing element (PE) array with inter-layer adaptive scheduling, eliminating resource waste caused by non-aligned layer dimensions; (2) it establishes a multi-level data reuse mechanism and a reconfiguration-free core array to maximize on-chip memory and computational resource utilization; (3) it presents the first TensorFlow-front-end-supported software-hardware co-compilation framework, enabling automatic tiling, dataflow scheduling, and RTL generation. Evaluated on the Xilinx VC709 platform, the architecture achieves a high sustained performance ratio relative to its theoretical peak, with significant improvements in throughput and energy efficiency. It further demonstrates high-frequency scalability and strong engineering practicality.
📝 Abstract
Though CNNs are highly parallel workloads, in the absence of efficient on-chip memory reuse techniques, an accelerator for them quickly becomes memory bound. In this paper, we propose a CNN accelerator design for inference that is able to exploit all forms of reuse available to minimize off-chip memory access while increasing utilization of available resources. The proposed design is composed of cores, each of which contains a one-dimensional array of processing elements. These cores can exploit different types of reuse available in CNN layers of varying shapes without requiring any reconfiguration; in particular, our design minimizes underutilization due to problem sizes that are not perfect multiples of the underlying hardware array dimensions. A major obstacle in the adoption of FPGAs as a platform for CNN inference is the difficulty to program these devices using hardware description languages. Our end goal is to also address this, and we develop preliminary software support via a codesign in order to leverage the accelerator through TensorFlow, a dominant high-level programming model. Our framework takes care of tiling and scheduling of neural network layers and generates necessary low-level commands to execute the CNN. Experimental evaluation on a real system with a PCI-express based Xilinx VC709 board demonstrates the effectiveness of our approach. As a result of an effective interconnection, the design maintains a high frequency when we scale the number of PEs. The sustained performance overall is a good fraction of the accelerator's theoretical peak performance.
Problem

Research questions and friction points this paper is trying to address.

Accelerating CNN inference on FPGAs efficiently
Minimizing off-chip memory access through reuse optimization
Overcoming FPGA programming difficulty via TensorFlow integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flexible FPGA accelerator for CNN inference
Exploits memory reuse to minimize off-chip access
TensorFlow integration with automatic tiling and scheduling
🔎 Similar Papers
2021-07-01IEEE International Conference on Application-Specific Systems, Architectures, and ProcessorsCitations: 13
Indian Institute of Science
K
Kingshuk Majumder
Dept of CSA, Indian Institute of Science, Bengaluru, India
U
Uday Bondhugula
Dept of CSA, Indian Institute of Science, Bengaluru, India