accelerator feature integration and runtime optimization

Designs and implements hardware–software integration layers, runtime systems, drivers, firmware, and compiler/back-end support that expose and use custom accelerator and silicon features. Builds and analyzes runtime scheduling, memory and resource management, and execution pipelines to optimize performance, latency, power efficiency, and interoperability between the silicon and higher-level frameworks.

acceleratorfeatureintegrationand

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$233K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Hardware and software build flow with SoCMake

Feb 04, 2025
RP
Risto Pejavsinovi'c
🏛️ CERN

ASIC development faces challenges in IP reuse and lacks integrated hardware-software co-verification and unified build infrastructure. Method: This paper introduces SoCMake—the first unified SoC build system supporting cross-compilation of Chisel/SystemRDL hardware descriptions with C/C++/assembly code. It integrates RTL generation, simulation, firmware compilation, and SoC configuration into a single workflow, overcoming the limited software compilation support of conventional hardware build tools. By deeply embedding SystemC, the RISC-V toolchain, and CMake’s extensibility framework, SoCMake enables automated, abstraction-level–aware co-building across hardware description → RTL → firmware. Contribution/Results: SoCMake has successfully accelerated iterative deployment of radiation-tolerant RISC-V SoCs in high-energy physics applications. After open-sourcing, it has become a de facto standard for generic SoC generation, reducing overall SoC development time by over 40% in empirical evaluations.

Addresses ASIC development cycle constraints.Automates fault-tolerant RISC-V SoC generation.Enhances hardware and software build system compatibility.

Understanding Accelerator Compilers via Performance Profiling

Nov 24, 2025
AY
Ayaka Yorihiro
🏛️ Cornell University

Accelerator Design Language (ADL) compilers suffer from unpredictable hardware performance due to the semantic gap between high-level abstractions and low-level implementations, compounded by reliance on heuristic optimizations. This work introduces Petal, the first tool enabling cycle-accurate, interpretable performance analysis for Calyx-based accelerators. Petal bridges the abstraction gap via three key techniques: source-code instrumentation, RTL-simulation trace collection, and a novel reverse-mapping algorithm that associates low-level timing events in synthesized hardware with high-level control-flow constructs. Crucially, it abandons the “compiler-perfection” assumption, instead exposing how concrete compilation decisions impact end-to-end latency. Evaluated on multiple real-world accelerator designs, Petal identifies subtle, manually elusive bottlenecks—enabling targeted manual optimizations that reduce total execution cycles by up to 46.9% for one application.

ADL compilers exhibit unpredictable performance due to complex optimizationsDevelopers lack tools to understand compiler decisions affecting performancePerformance problems in generated hardware require guesswork to resolve

This work addresses the limitations of traditional application-specific hardware accelerators, which suffer from large area overhead and low utilization, as well as the inability of existing reconfigurable processors to support microcode-level dynamic control flow—such as loops, conditional branches, and exception handling—hindering their efficiency on compute-intensive tasks with complex control logic. To overcome these challenges, this paper introduces, for the first time, a complete dynamic control flow execution mechanism at the microcode level of a runtime-reconfigurable processor. This enables flexible switching of accelerator configurations during execution and facilitates efficient collaboration between general-purpose cores and configurable accelerators. The proposed approach significantly enhances system flexibility and applicability, achieving substantial speedups over conventional general-purpose processors in diverse domains including object detection, ocean simulation, artificial intelligence, and security.

compute-intensive applicationsdynamic control-flowhardware accelerators

To address high power consumption, inflexible instruction sets, and difficulties in integrating domain-specific accelerators in edge-computing embedded systems, this work designs and implements a heterogeneous SoC based on the RISC-V RV32I+M+A ISA, incorporating a tightly coupled, custom DSP accelerator. Leveraging RISC-V’s modular ISA, we propose a software–hardware co-optimization architecture that enables instruction-level and microarchitectural-level coordination while preserving full standard compliance. The design employs cycle-accurate simulation and RTL-level integration, combined with low-power circuit techniques. Under identical process technology, it achieves a 17% reduction in dynamic power versus the ARM Cortex-M0 and a significant reduction in CPI. Our key contribution is the first lightweight, tightly coupled accelerator microarchitecture specifically tailored for edge-oriented DSP workloads—demonstrating, for the first time, simultaneous improvements in energy efficiency and real-time performance for RISC-V-based heterogeneous SoCs.

Comparing RISC-V power consumption with ARM Cortex-M0 implementationsDesigning a RISC-V SoC with custom DSP accelerators for edge computingEvaluating performance and power efficiency of RISC-V ISA extensions

This work addresses the inefficiency of existing programmable architectures in handling sparse or irregular data and the inflexibility of dedicated accelerators when confronted with new kernels or input patterns. To bridge this gap, the paper proposes Canon, a novel architecture that integrates a programmable finite state machine (FSM) with a dynamic, data-driven execution orchestration mechanism to generate control flow at runtime. Canon further introduces a time-interleaved SIMD execution model that constructs an evolving dataflow to maximize parallelism. This design achieves performance and energy efficiency approaching that of specialized accelerators across a range of data-oblivious and data-driven kernels, while preserving the programmability and flexibility of general-purpose architectures.

execution orchestrationirregular dataperformance fragility

Latest Papers

What's happening recently
View more

Traditional approaches struggle to provide deep visibility into the internal behavior of the gem5 simulator. This work proposes a non-intrusive, lightweight runtime call-stack analysis framework that, for the first time, treats the simulator’s own execution path as a novel lens for understanding simulated system behavior. Built upon the Linux perf_event interface, the framework enables parallel sampling, real-time symbol resolution, and hierarchical call-tree aggregation, with support for component-level customizable analysis. Experimental results demonstrate its effectiveness in uncovering performance bottlenecks in TimingSimpleCPU and identifying deadlock and livelock issues within the Ruby memory system—capturing critical behavioral characteristics that conventional statistical methods fail to detect.

cache coherencecall-stack profilinggem5

Deploying GEMM on tile-based multi-PE accelerators faces challenges of deployment complexity and deep hardware-software coupling. To address this, we propose an end-to-end automated deployment framework. Our approach introduces the novel “Design in Tiles” paradigm, integrating configurable execution modeling, hardware-aware automatic mapping, hierarchical tiling scheduling, and compute-memory co-optimized compilation. For the first time, we achieve superior PE utilization over NVIDIA GH200’s expert-tuned library on a large-scale 32×32 tile configuration. At FP8 precision, our framework delivers 1979 TFLOPS peak performance and accelerates diverse matrix shapes by 1.2–2.0× relative to GH200. This work bridges the compilation gap between configurable hardware architectures and high-level computational graphs, establishing a general, efficient, and scalable methodology for automatic mapping onto domain-specific accelerators.

Addresses programming difficulty due to hardware-software couplingAutomates GEMM deployment on tile-based many-PE acceleratorsImproves performance over expert-tuned libraries on large configurations

This work proposes the first modular profiling framework tailored for hardware accelerators, addressing the lack of low-overhead and flexible program analysis tools in modern computing systems. By abstracting underlying performance APIs and integrating with mainstream deep learning frameworks, the framework offers a unified interface to capture runtime events across multiple abstraction levels and enables rapid prototyping. It features a GPU-accelerated backend, multi-level event tracing, and cross-platform compatibility (NVIDIA/AMD), achieving high scalability and minimal profiling overhead in both single- and multi-GPU settings. Experimental results demonstrate that, on representative deep learning workloads, the framework achieves up to 1.3×10⁴ times faster profiling compared to conventional tools while delivering fine-grained performance insights.

hardware acceleratorslow-overheadmodular framework

This work addresses the inefficiencies in hardware-software co-integration of modern accelerators, which stem from architectural complexity, deep memory hierarchies, and heavy reliance on production firmware. Traditional FPGA-based simulation workflows suffer from slow debugging cycles and prolonged iteration times. To overcome these limitations, we present the first framework enabling cycle-accurate co-verification of production firmware with RTL or gate-level hardware models within standard simulators such as VCS, Xcelium, and Vivado Xsim. By compiling firmware to x86 and bridging it with the hardware emulation subsystem—augmented with a randomized memory bridge—the framework supports second-scale debugging, register-level protocol validation, off-chip dataflow analysis, and memory congestion emulation. Evaluated on accelerators including systolic arrays and CGRAs, our approach achieves up to 50× faster debugging and significantly enhances parallel development efficiency and functional verification reliability for heterogeneous computing platforms.

accelerator integrationcycle-accurate simulationdebug iteration

Hot Scholars

NC

Ningyuan Cao

University of Notre Dame
Hardware for machine learningIC design automation
SB

Sven Beyer

Globalfoundries
CMOSFeFETferroelectriceNVM
SY

Shimeng Yu

Georgia Institute of Technology, Dean's Professor
Non-volatile MemoryRRAMFerroelectric MemoriesIn-Memory Computing
AS

Abhronil Sengupta

Monkowski Career Development Associate Professor of EECS, Penn State University
Neuromorphic Computing