🤖 AI Summary
This study addresses the acceleration challenges of non-matrix machine learning kernels regarding low latency and high energy efficiency by proposing a spatio-temporally elastic coarse-grained reconfigurable array (CGRA) architecture fabricated in a 16nm process. The design achieves efficient spatial computing through fast configuration paths, single-PE hardware loop control, and a deeply pipelined elastic interconnect. Furthermore, it introduces a spatio-temporal data reuse mechanism that overcomes traditional matrix-centric limitations, significantly optimizing performance for multi-stage pipelined applications. Experimental results demonstrate that the proposed architecture attains an area efficiency of 110 GOPS/mm² at 850 MHz, improving effective kernel throughput by 4.84× to 7.14× compared to the baseline.
📝 Abstract
Emerging non-matrix ML kernels, such as LayerNorm, GeLu, FFT or circular convolutions, demand low-latency, energy-efficient spatial accelerators beyond MatMul-centric arrays. STELLA presents a spatio-temporal elastic 16 nm coarse-grained reconfigurable array (CGRA) with a rapid configuration path, per-PE hardware loop control, and a low-latency, deeply pipelined elastic fabric with spatio-temporal data reuse. STELLA reaches up to 110 GOPS/mm2 at 850 MHz, and improves effective kernel throughput by 4.84-7.14x over baseline CGRAs.