STELLA: A 16nm Spatio-Temporal Elastic Low-Latency CGRA for Multi-Stage Pipelined Applications

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the acceleration challenges of non-matrix machine learning kernels regarding low latency and high energy efficiency by proposing a spatio-temporally elastic coarse-grained reconfigurable array (CGRA) architecture fabricated in a 16nm process. The design achieves efficient spatial computing through fast configuration paths, single-PE hardware loop control, and a deeply pipelined elastic interconnect. Furthermore, it introduces a spatio-temporal data reuse mechanism that overcomes traditional matrix-centric limitations, significantly optimizing performance for multi-stage pipelined applications. Experimental results demonstrate that the proposed architecture attains an area efficiency of 110 GOPS/mm² at 850 MHz, improving effective kernel throughput by 4.84× to 7.14× compared to the baseline.
📝 Abstract
Emerging non-matrix ML kernels, such as LayerNorm, GeLu, FFT or circular convolutions, demand low-latency, energy-efficient spatial accelerators beyond MatMul-centric arrays. STELLA presents a spatio-temporal elastic 16 nm coarse-grained reconfigurable array (CGRA) with a rapid configuration path, per-PE hardware loop control, and a low-latency, deeply pipelined elastic fabric with spatio-temporal data reuse. STELLA reaches up to 110 GOPS/mm2 at 850 MHz, and improves effective kernel throughput by 4.84-7.14x over baseline CGRAs.
Problem

Research questions and friction points this paper is trying to address.

non-matrix ML kernels
low-latency
energy-efficient
spatial accelerators
CGRA
Innovation

Methods, ideas, or system contributions that make the work stand out.

CGRA
spatio-temporal elasticity
low-latency pipeline
non-matrix ML kernels
hardware loop control
🔎 Similar Papers
No similar papers found.