KWT-Tiny: RISC-V Accelerated, Embedded Keyword Spotting Transformer

πŸ“… 2024-07-22
πŸ›οΈ ACM Symposium on Cloud Computing
πŸ“ˆ Citations: 1
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Resource-constrained IoT devices face severe memory and computational bottlenecks for deploying Transformer-based keyword spotting (KWS) models at the edge. Method: This work proposes a hardware-software co-optimization framework for RISC-V architectures, integrating structured pruning with retraining, INT8 quantization, a custom lightweight edge AI library, andβ€”novellyβ€”RISC-V custom instructions to accelerate GELU and Softmax operations. Contribution/Results: The optimized Keyword Transformer achieves a 369Γ— model compression (2.42 MB β†’ 1.65 KB), operates in bare-metal C with only 64 KB RAM, reduces inference latency to 5.5M cycles (5Γ— speedup), cuts energy consumption by ~5Γ—, and incurs only a 10% accuracy drop. To our knowledge, this is the first demonstration of a Transformer-based KWS model operating efficiently under ultra-low memory constraints (1.65 KB), establishing a deployable paradigm for RISC-V-based edge AI.

Technology Category

Machine Learning: Hardware-aware MLComputer Vision: Learning & Optimization for CVSearch and Optimization: Learning to Search

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsResponsible Web: Data and user privacy-enhancing technologies for the WebSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search engines
πŸ“ Abstract
This paper explores the adaptation of Transformer-based models for edge devices through the quantisation and hardware acceleration of the ARM Keyword Transformer (KWT) model on a RISC-V platform. The model was targeted to run on 64kB RAM in bare-metal C using a custom-developed edge AI library. KWT-1 was retrained to be 369 times smaller, with only a 10% loss in accuracy through reducing output classes from 35 to 2. The retraining and quantisation reduced model size from 2.42 MB to 1.65 kB. The integration of custom RISC-V instructions that accelerated GELU and SoftMax operations enabled a 5x speedup and thus ~5x power reduction in inference, with inference clock cycle counts decreasing from 26 million to 5.5 million clock cycles while incurring a small area overhead of approximately 29%. The results demonstrate a viable method for porting and accelerating Transformer-based models in low-power IoT devices.
Problem

Research questions and friction points this paper is trying to address.

Adapting Transformer models for edge devices with limited resources
Reducing model size and power consumption for IoT applications
Accelerating keyword spotting on RISC-V platform through hardware optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantised Transformer model for edge devices
Custom RISC-V instructions accelerate operations
Retrained model reduced size with minimal accuracy loss
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
University College Dublin
A
Aness Al-Qawlaq
University College Dublin, Ireland
A
Ajay Kumar
University College Dublin, Ireland
Deepu John
Deepu John
University College Dublin
Edge ComputingIoTWearable SensingBiomedical Circuits and Systems