π€ AI Summary
Resource-constrained IoT devices face severe memory and computational bottlenecks for deploying Transformer-based keyword spotting (KWS) models at the edge.
Method: This work proposes a hardware-software co-optimization framework for RISC-V architectures, integrating structured pruning with retraining, INT8 quantization, a custom lightweight edge AI library, andβnovellyβRISC-V custom instructions to accelerate GELU and Softmax operations.
Contribution/Results: The optimized Keyword Transformer achieves a 369Γ model compression (2.42 MB β 1.65 KB), operates in bare-metal C with only 64 KB RAM, reduces inference latency to 5.5M cycles (5Γ speedup), cuts energy consumption by ~5Γ, and incurs only a 10% accuracy drop. To our knowledge, this is the first demonstration of a Transformer-based KWS model operating efficiently under ultra-low memory constraints (1.65 KB), establishing a deployable paradigm for RISC-V-based edge AI.
π Abstract
This paper explores the adaptation of Transformer-based models for edge devices through the quantisation and hardware acceleration of the ARM Keyword Transformer (KWT) model on a RISC-V platform. The model was targeted to run on 64kB RAM in bare-metal C using a custom-developed edge AI library. KWT-1 was retrained to be 369 times smaller, with only a 10% loss in accuracy through reducing output classes from 35 to 2. The retraining and quantisation reduced model size from 2.42 MB to 1.65 kB. The integration of custom RISC-V instructions that accelerated GELU and SoftMax operations enabled a 5x speedup and thus ~5x power reduction in inference, with inference clock cycle counts decreasing from 26 million to 5.5 million clock cycles while incurring a small area overhead of approximately 29%. The results demonstrate a viable method for porting and accelerating Transformer-based models in low-power IoT devices.