Score
Designs, builds, and analyzes optimized software and system configurations for resource‑constrained embedded devices, implementing hardware‑aware and on‑device code optimizations (including SIMD and runtime‑level tuning), RPC and system‑level performance improvements, and real‑time scheduling to meet latency, jitter, and timing constraints. Uses profiling and benchmarking on target hardware and system‑level tradeoff analysis to identify and tune bottlenecks and validate that performance, latency, and control‑signal quality requirements are satisfied.
Real-time optimal control (e.g., model predictive control) for resource-constrained robots faces hardware bottlenecks in latency, energy efficiency, and computational density. Method: This paper proposes a hardware–software co-design methodology for embedded SoC architecture space exploration, centered on robotic control workloads. It introduces a quantitative evaluation framework integrating kernel-level benchmarks with end-to-end task-driven analysis (e.g., motion control, manipulation), coupled with hardware modeling, software stack optimization, and a custom compiler-enabled automated code generation pipeline. Contribution/Results: The work systematically compares scalar CPUs, vector processors, and domain-specific accelerators across performance, area, and utilization trade-offs—revealing that dedicated accelerators reduce control latency and improve energy efficiency by over 3.2× versus general-purpose processors. Furthermore, the developed mapping and code-generation toolchain demonstrates strong reusability across robotic control applications.
This work addresses the limitation of conventional hardware compilation flows, wherein pipeline optimization is deferred to the backend, resulting in the loss of high-level structural information and suboptimal global optimization. To overcome this, the paper introduces, for the first time, an explicit pipeline-aware compiler pass at the intermediate representation (IR) level. By modeling legality constraints for register relocation, integrating learning-driven timing prediction, and formulating the timing-constrained relocation problem as a global minimum-cost flow problem, the approach enables timing-aware RTL generation. Implemented within the CIRCT framework, the method significantly reduces critical-path delay, power, and area across both open-source and commercial designs, while also providing backend retiming with a superior initial structure.
Accelerator Design Language (ADL) compilers suffer from unpredictable hardware performance due to the semantic gap between high-level abstractions and low-level implementations, compounded by reliance on heuristic optimizations. This work introduces Petal, the first tool enabling cycle-accurate, interpretable performance analysis for Calyx-based accelerators. Petal bridges the abstraction gap via three key techniques: source-code instrumentation, RTL-simulation trace collection, and a novel reverse-mapping algorithm that associates low-level timing events in synthesized hardware with high-level control-flow constructs. Crucially, it abandons the “compiler-perfection” assumption, instead exposing how concrete compilation decisions impact end-to-end latency. Evaluated on multiple real-world accelerator designs, Petal identifies subtle, manually elusive bottlenecks—enabling targeted manual optimizations that reduce total execution cycles by up to 46.9% for one application.
RISC-V platforms suffer from unreliable performance analysis due to toolchain fragmentation, immature hardware performance monitoring units (PMUs), and architectural limitations. To address this, we propose a compiler-driven, hardware-agnostic Roofline modeling methodology. Our approach leverages LLVM IR-level instrumentation and software-based compensation to bypass defective hardware counters, enabling robust event sampling. We innovatively employ compiler static analysis to derive operational intensity and throughput metrics—eliminating dependence on PMU measurements. Furthermore, we design an automated workflow that jointly calibrates PMU data, constructs Roofline models, and validates their accuracy. Evaluated across multiple emerging RISC-V chips, our method demonstrates effectiveness in identifying performance bottlenecks. The open-source toolchain supports cross-platform analysis and significantly improves the reliability of performance assessment and optimization efficiency—especially in scenarios where mature PMUs are unavailable.
In the post-Dennard era, embedded systems face intricate trade-offs between energy efficiency and latency, rendering traditional heuristic methods ineffective in navigating the high-dimensional, non-smooth scheduling space. This work proposes a Gaussian process-based multi-objective Bayesian optimization framework to automatically discover Pareto-optimal scheduling strategies that balance energy consumption and execution time on heterogeneous multicore architectures. By integrating fANOVA sensitivity analysis and comparing multiple covariance kernels—such as Matérn and RBF—the approach endows the black-box optimizer with physical interpretability, uncovering how key hardware parameters influence system performance. Experimental results demonstrate that the method efficiently approximates the Pareto front, significantly advancing both the automation of scheduling and the understanding of underlying hardware behaviors.
Existing heterogeneous accelerator designs primarily prioritize throughput or quality of service, falling short in meeting the stringent requirements of safety-critical real-time systems—namely, predictability, real-time awareness, and rigorous schedulability. This work proposes PHAROS, a novel framework that, for the first time, integrates modern real-time scheduling theory into heterogeneous accelerator design. PHAROS introduces a preemptive scheduling mechanism supporting both FIFO and EDF policies and formulates a soft real-time schedulability analysis model. Building upon this foundation, it develops a schedulability-driven design space exploration algorithm. Experimental results demonstrate that PHAROS significantly improves task set schedulability and real-time responsiveness across diverse applications, uncovering a substantially broader range of feasible configurations compared to throughput-oriented approaches.
Existing benchmarks struggle to evaluate the end-to-end capabilities of large language models (LLMs) in system-level hardware-software co-design, often assessing hardware and software components in isolation. This work introduces the first benchmark that encompasses the full co-design workflow, requiring an LLM agent to analyze applications, design heterogeneous accelerators, map kernel functions, and deploy a complete system-on-chip (SoC) prototype on an AMD VC707 FPGA. Built upon an open-source SoC platform and a structured repository, the benchmark enables LLMs to jointly reason about and modify both hardware and software stacks. Experimental results show that among five state-of-the-art models, only two successfully generated functional prototypes, achieving a peak speedup of 16.22×, yet with a maximum resource utilization of merely 23.67%, indicating that current LLMs have not yet fully harnessed the potential of hardware acceleration.
This study addresses the immature compiler support for automatic vectorization on real hardware implementing the RISC-V Vector Extension (RVV 1.0), which limits its performance in scientific computing and machine learning. We present the first systematic evaluation of automatic vectorization capabilities in GCC 15 and LLVM 21 on RVV hardware, combining assembly-level microbenchmarks, perf counter calibration, and comparative experiments between manual and compiler-generated vectorization using the Qsim quantum simulator. Our analysis reveals that key performance bottlenecks—such as predicate overhead and strided memory accesses—are inadequately modeled by current cost models, while default LMUL selection is already near-optimal. Experimental results show GCC 15 outperforms LLVM 21 in four of six proxy applications; LLVM’s advantage in SGEMM/DGEMM stems from aggressive instruction reduction, highlighting both compilers’ insufficient handling of complex memory access patterns.
Traditional performance simulators are slow and costly, while existing machine learning approaches suffer from limitations in accuracy, speed, or coverage. This work proposes a hierarchical LSTM-based AI model that leverages microarchitecture-agnostic program execution feature traces to efficiently predict full-benchmark performance metrics without requiring detailed simulation or instruction-level encoding. The method achieves, for the first time, high-accuracy, end-to-end whole-program performance prediction, attaining an average IPC prediction error of only 9.35% on SPEC CPU 2017. The entire prediction process completes in just 2 minutes and 57 seconds—offering accuracy comparable to state-of-the-art techniques while accelerating prediction by three orders of magnitude.
This work addresses the challenges of fragmented and heterogeneous performance monitoring units in modern heterogeneous multi-core SoCs, which complicate data collection, synchronization, and correlation. The paper proposes a centralized performance monitoring architecture—demonstrated for the first time in a RISC-V SoC—that enables unified cross-component monitoring. Hardware microarchitectural events are collected by Event Monitoring Units (EVUs) and routed via the AXI4 bus to an Advanced Performance Monitoring Unit (APMU) for consolidated processing. By integrating programmable counters and a dedicated processing unit, the architecture simplifies interface design, enhances event correlation capabilities, and supports event-driven software mechanisms. Experimental results validate its effectiveness in real-time resource management, application profiling, and counter attribution scenarios.