🤖 AI Summary
This study addresses the inefficient Transformer inference on resource-constrained edge AI devices. To overcome this limitation, it proposes a hardware-software co-optimization framework tailored for the RISC-V architecture. On the hardware side, custom instruction set extensions are designed to accelerate address generation for matrix multiplication. On the software side, compiler-based loop unrolling strategies are integrated to facilitate the efficient deployment of lightweight language models on Internet of Things (IoT) devices. Validated through ASIC and FPGA implementations, the proposed approach achieves up to a 2.19× improvement in inference speed with minimal hardware overhead, increasing area and power consumption by only 21% and 2.33%, respectively. This work provides a highly energy-efficient solution for deploying large models at the edge.
📝 Abstract
This work presents a framework for accelerating transformer-based language models (LMs) on resource-constrained IoT devices. The framework targets compact LMs: BERT-Tiny (B-Ty), MobileBERT (M-Bt), MiniLM (M-Lm), Electra (E-Lt) and DeBERTa (D-Bt) -- selected for their architectural diversity and use in edge inference scenarios. The proposed flow derives lightweight instruction set extensions tailored to the non-obvious computational patterns of these models. In addition, a custom instruction is introduced to accelerate the address generation stage of batch matrix multiplication, achieving a performance improvement of 15.39--21.74% with modest ASIC overheads of 6.79% in area and 2.33% in power. To further enhance performance without incurring additional processor core hardware cost, an optional compiler-directed loop unrolling strategy is employed, trading increased code size for overall reduced execution time. Evaluation on the Synopsys trv32p3f RISC-V core demonstrates inference speedups of up to 2.19x, and FPGA implementation on the AMD Zynq UltraScale+ ZCU102 shows a 32.94% area overhead at 75 MHz, whereas the ASIC implementation using the TSMC 28 nm library incurs a 21.08% area overhead while operating at 250 MHz.