🤖 AI Summary
Existing DNN accelerators struggle to balance energy efficiency and programmability, limiting their adaptability to rapidly evolving models. This work proposes a RISC-V-inspired custom instruction set architecture coupled with a reconfigurable hardware platform that decouples execution from data movement, enables fine-grained control over data transfers, and supports dynamic precision computation from 3 to 8 bits using the residue number system (RNS), thereby achieving representation-agnostic flexibility. By integrating lightweight programmable cores with reconfigurable SIMD units, the architecture achieves an energy efficiency of 5.12–10.47 TOPS/W on typical workloads in a 22nm process—up to 1.2× higher than fixed-point implementations and surpassing state-of-the-art mixed-precision accelerators—while preserving model accuracy.
📝 Abstract
Domain-specific hardware accelerators provide significantly higher performance and energy efficiency for deep neural network (DNN) workloads than general-purpose processors, but often lack adaptability to evolving model architectures. In contrast, general-purpose ISA-based solutions, such as RISC-V-based accelerators, improve programmability at the cost of efficiency. This work addresses this tradeoff by introducing a machine-learning-oriented instruction set architecture (ISA) and a reconfigurable hardware platform that combine high efficiency with flexibility. The proposed ISA enables fine-grained control over data movement, dynamic precision, and decoupled execution across data-fetching, tensor processing, and post-processing domains. The corresponding architecture employs lightweight programmable cores and SIMD units to maintain high processing-element utilization with low control overhead, while remaining independent of the underlying numerical representation. We demonstrate the approach using a Residue Number System (RNS) instantiation supporting 3-8-bit dynamic precision. A 22-nm implementation achieves 5.12-10.47 TOPS/W for a typical workload and up to 1.2x higher energy efficiency than its fixed-point counterpart, while preserving model accuracy. It also outperforms state-of-the-art and mixed-precision accelerators. These results show that the proposed design effectively bridges the gap between efficiency and programmability in modern DNN accelerators.