🤖 AI Summary
This study addresses the storage redundancy and lack of runtime adaptability caused by requiring multiple binaries for different pruning rates when deploying Vision Transformers (ViTs) on edge devices. To overcome these limitations, this work proposes an end-to-end deployment pipeline that converts pretrained ViTs into a single configurable binary supporting dynamic sparsity switching. Methodologically, it introduces a hardware-aligned MLP block pruning strategy with C-kernel loop restructuring, alongside customized RISC-V ISA extensions exploiting input reuse patterns in linear projections. Implementation using a TSMC 28nm process demonstrates a 4.86× reduction in storage overhead and a 2.8× inference speedup. Furthermore, the proposed ISA extensions yield an additional 1.56× acceleration and a 33% decrease in energy consumption, incurring only a 24.7% area overhead.
📝 Abstract
Deploying Vision Transformers (ViTs) on low-power edge devices is challenging due to high computational demands. Conventional pruning frameworks require a separate compiled binary for each sparsity level, increasing storage overhead and limiting runtime adaptability. This paper presents an end-to-end deployment pipeline that transforms pretrained ViTs into a single runtime-configurable binary, enabling dynamic compute-budget switching on embedded CPUs. This is achieved by restructuring generated C kernels with modified loop bounds and binary-mask control logic, allowing execution to switch across discrete sparsity levels via compact external configuration files. Compared to multi-binary deployment, the proposed runtime-adaptive approach reduces on-device storage by up to 4.86x, requiring only 163 MB for ViT-Base instead of nearly 800 MB. To maximize pruning efficiency, we introduce a hardware-aligned block pruning strategy for Multi-Layer Perceptron (MLP) layers. In addition, a custom ISA extension is proposed to exploit input-reuse patterns in linear projection kernels. On a Synopsys TRV32P3FX RISC-V processor, the full system achieves up to 2.8x speedup at 65% MLP and 50% attention-head pruning for ViT-Base. The ISA extension alone provides a 1.56x speedup and 33% lower inference energy, with a 24.7% area overhead in a TSMC 28 nm implementation.