Score
Designs and implements hardware–software integration layers, runtime systems, drivers, firmware, and compiler/back-end support that expose and use custom accelerator and silicon features. Builds and analyzes runtime scheduling, memory and resource management, and execution pipelines to optimize performance, latency, power efficiency, and interoperability between the silicon and higher-level frameworks.
ASIC development faces challenges in IP reuse and lacks integrated hardware-software co-verification and unified build infrastructure. Method: This paper introduces SoCMake—the first unified SoC build system supporting cross-compilation of Chisel/SystemRDL hardware descriptions with C/C++/assembly code. It integrates RTL generation, simulation, firmware compilation, and SoC configuration into a single workflow, overcoming the limited software compilation support of conventional hardware build tools. By deeply embedding SystemC, the RISC-V toolchain, and CMake’s extensibility framework, SoCMake enables automated, abstraction-level–aware co-building across hardware description → RTL → firmware. Contribution/Results: SoCMake has successfully accelerated iterative deployment of radiation-tolerant RISC-V SoCs in high-energy physics applications. After open-sourcing, it has become a de facto standard for generic SoC generation, reducing overall SoC development time by over 40% in empirical evaluations.
Accelerator Design Language (ADL) compilers suffer from unpredictable hardware performance due to the semantic gap between high-level abstractions and low-level implementations, compounded by reliance on heuristic optimizations. This work introduces Petal, the first tool enabling cycle-accurate, interpretable performance analysis for Calyx-based accelerators. Petal bridges the abstraction gap via three key techniques: source-code instrumentation, RTL-simulation trace collection, and a novel reverse-mapping algorithm that associates low-level timing events in synthesized hardware with high-level control-flow constructs. Crucially, it abandons the “compiler-perfection” assumption, instead exposing how concrete compilation decisions impact end-to-end latency. Evaluated on multiple real-world accelerator designs, Petal identifies subtle, manually elusive bottlenecks—enabling targeted manual optimizations that reduce total execution cycles by up to 46.9% for one application.
This work addresses the limitations of traditional application-specific hardware accelerators, which suffer from large area overhead and low utilization, as well as the inability of existing reconfigurable processors to support microcode-level dynamic control flow—such as loops, conditional branches, and exception handling—hindering their efficiency on compute-intensive tasks with complex control logic. To overcome these challenges, this paper introduces, for the first time, a complete dynamic control flow execution mechanism at the microcode level of a runtime-reconfigurable processor. This enables flexible switching of accelerator configurations during execution and facilitates efficient collaboration between general-purpose cores and configurable accelerators. The proposed approach significantly enhances system flexibility and applicability, achieving substantial speedups over conventional general-purpose processors in diverse domains including object detection, ocean simulation, artificial intelligence, and security.
To address high power consumption, inflexible instruction sets, and difficulties in integrating domain-specific accelerators in edge-computing embedded systems, this work designs and implements a heterogeneous SoC based on the RISC-V RV32I+M+A ISA, incorporating a tightly coupled, custom DSP accelerator. Leveraging RISC-V’s modular ISA, we propose a software–hardware co-optimization architecture that enables instruction-level and microarchitectural-level coordination while preserving full standard compliance. The design employs cycle-accurate simulation and RTL-level integration, combined with low-power circuit techniques. Under identical process technology, it achieves a 17% reduction in dynamic power versus the ARM Cortex-M0 and a significant reduction in CPI. Our key contribution is the first lightweight, tightly coupled accelerator microarchitecture specifically tailored for edge-oriented DSP workloads—demonstrating, for the first time, simultaneous improvements in energy efficiency and real-time performance for RISC-V-based heterogeneous SoCs.
This work addresses the inefficiency of existing programmable architectures in handling sparse or irregular data and the inflexibility of dedicated accelerators when confronted with new kernels or input patterns. To bridge this gap, the paper proposes Canon, a novel architecture that integrates a programmable finite state machine (FSM) with a dynamic, data-driven execution orchestration mechanism to generate control flow at runtime. Canon further introduces a time-interleaved SIMD execution model that constructs an evolving dataflow to maximize parallelism. This design achieves performance and energy efficiency approaching that of specialized accelerators across a range of data-oblivious and data-driven kernels, while preserving the programmability and flexibility of general-purpose architectures.
Traditional approaches struggle to provide deep visibility into the internal behavior of the gem5 simulator. This work proposes a non-intrusive, lightweight runtime call-stack analysis framework that, for the first time, treats the simulator’s own execution path as a novel lens for understanding simulated system behavior. Built upon the Linux perf_event interface, the framework enables parallel sampling, real-time symbol resolution, and hierarchical call-tree aggregation, with support for component-level customizable analysis. Experimental results demonstrate its effectiveness in uncovering performance bottlenecks in TimingSimpleCPU and identifying deadlock and livelock issues within the Ruby memory system—capturing critical behavioral characteristics that conventional statistical methods fail to detect.
Deploying GEMM on tile-based multi-PE accelerators faces challenges of deployment complexity and deep hardware-software coupling. To address this, we propose an end-to-end automated deployment framework. Our approach introduces the novel “Design in Tiles” paradigm, integrating configurable execution modeling, hardware-aware automatic mapping, hierarchical tiling scheduling, and compute-memory co-optimized compilation. For the first time, we achieve superior PE utilization over NVIDIA GH200’s expert-tuned library on a large-scale 32×32 tile configuration. At FP8 precision, our framework delivers 1979 TFLOPS peak performance and accelerates diverse matrix shapes by 1.2–2.0× relative to GH200. This work bridges the compilation gap between configurable hardware architectures and high-level computational graphs, establishing a general, efficient, and scalable methodology for automatic mapping onto domain-specific accelerators.
This work proposes the first modular profiling framework tailored for hardware accelerators, addressing the lack of low-overhead and flexible program analysis tools in modern computing systems. By abstracting underlying performance APIs and integrating with mainstream deep learning frameworks, the framework offers a unified interface to capture runtime events across multiple abstraction levels and enables rapid prototyping. It features a GPU-accelerated backend, multi-level event tracing, and cross-platform compatibility (NVIDIA/AMD), achieving high scalability and minimal profiling overhead in both single- and multi-GPU settings. Experimental results demonstrate that, on representative deep learning workloads, the framework achieves up to 1.3×10⁴ times faster profiling compared to conventional tools while delivering fine-grained performance insights.
This work addresses the inefficiencies in hardware-software co-integration of modern accelerators, which stem from architectural complexity, deep memory hierarchies, and heavy reliance on production firmware. Traditional FPGA-based simulation workflows suffer from slow debugging cycles and prolonged iteration times. To overcome these limitations, we present the first framework enabling cycle-accurate co-verification of production firmware with RTL or gate-level hardware models within standard simulators such as VCS, Xcelium, and Vivado Xsim. By compiling firmware to x86 and bridging it with the hardware emulation subsystem—augmented with a randomized memory bridge—the framework supports second-scale debugging, register-level protocol validation, off-chip dataflow analysis, and memory congestion emulation. Evaluated on accelerators including systolic arrays and CGRAs, our approach achieves up to 50× faster debugging and significantly enhances parallel development efficiency and functional verification reliability for heterogeneous computing platforms.