🤖 AI Summary
This study addresses the challenges of full-stack integration for heterogeneous accelerators and the substantial overhead associated with hardware-software co-development. We propose an Extensible Accelerator Architecture based on MLIR (EAAC), which employs RISC-V as the host core and systolic arrays as acceleration units. By providing minimal data orchestration capabilities and an unbiased design starting point, EAAC effectively lowers the barrier to hardware-compiler co-development, streamlines the implementation of acceleration units, and supports rapid prototyping. Experimental results demonstrate that the EAAC-based GEMM accelerator achieves a 28× speedup over a pure software baseline, with computation accounting for only 1.64% of the total execution time. These findings indicate that the proposed architecture significantly enhances both development efficiency and execution performance for heterogeneous systems.
📝 Abstract
Heterogeneous accelerator architectures offer an efficient path to performance for compute-intensive workloads. However, full-stack integration remains difficult. We present EAAC (Extensible Accelerator Architecture), a flexible and extensible compiler and hardware architecture designed to lower the overhead of hardware-compiler co-development for rapid prototyping of hardware accelerators. EAAC targets static data-flow workloads with predictable memory access patterns.
By using MLIR, we enable possible integration with a range of different frontends that emit MLIR. And by providing a minimal set of compiler functionality that enable necessary data-orchestration we lower the effort needed to get a simple implementation of a hardware acceleration unit up and running, while providing a relatively blank and un-opinionated starting point for further work.
We validate this approach with a prototype GEMM accelerator combining a systolic array unit with an embedded RISC-V core, compiled end-to-end through the EAAC MLIR pipeline. On a synthetic fully-connected-layer workload, the accelerator achieves a 28x speedup in execution time over a RISC-V-only baseline, with the GEMM operation itself accounting for only 1.64\% of total execution time. We further characterize the compiler's hardware-semaphore allocation, showing that the number of semaphores required scales linearly with instruction count in the worst case, and identify this as a concrete target for future optimization.