🤖 AI Summary
This study addresses the limited end-to-end support for TensorFlow Lite INT8 inference in existing open-source accelerators, which hinders the efficient deployment of low-precision CNNs in edge AI. To this end, we propose ZTA-Q, an open-source platform built upon the RISC-V architecture that achieves accurate deployment of TFLite INT8 models by extending quantized operators and introducing a novel configurable post-processing datapath. Furthermore, this work systematically investigates the impact of circuit-level approximations—including multiplier precision, shift-scaling, and rounding simplifications—on model accuracy. Experimental results demonstrate that ZTA-Q operates at 83.3 MHz on an Arty A7-100T FPGA while strictly constraining Top-1/Top-5 accuracy degradation to within 0.25 percentage points. Ultimately, this platform provides a comprehensive hardware-software co-design solution for efficient and accurate quantized inference at the edge.
📝 Abstract
Low-precision inference is widely adopted in edge AI to reduce computational cost and memory footprint. However, existing open-source accelerator platforms provide limited end-to-end support for CNNs following the standard TensorFlow Lite integer inference scheme. This paper presents ZTA-Q, an open-source RISC-V-based platform that enables accurate deployment of TensorFlow Lite INT8 models. In addition to extending operator support, ZTA-Q provides a configurable post-processing datapath for studying how circuit-level approximations, including reduced multiplier precision, shared shift scaling, and simplified rounding, affect model accuracy. The proposed system is implemented on a Digilent Arty A7-100T FPGA and operates at 83.3 MHz. Evaluations on representative CNN models show that with LUT, register, and DSP overheads of 26.3%, 12.6%, and 150%, respectively, ZTA-Q limits the degradation in both top-1 and top-5 accuracy to within 0.25 percentage points.