🤖 AI Summary
RISC-V platforms suffer from unreliable performance analysis due to toolchain fragmentation, immature hardware performance monitoring units (PMUs), and architectural limitations. To address this, we propose a compiler-driven, hardware-agnostic Roofline modeling methodology. Our approach leverages LLVM IR-level instrumentation and software-based compensation to bypass defective hardware counters, enabling robust event sampling. We innovatively employ compiler static analysis to derive operational intensity and throughput metrics—eliminating dependence on PMU measurements. Furthermore, we design an automated workflow that jointly calibrates PMU data, constructs Roofline models, and validates their accuracy. Evaluated across multiple emerging RISC-V chips, our method demonstrates effectiveness in identifying performance bottlenecks. The open-source toolchain supports cross-platform analysis and significantly improves the reliability of performance assessment and optimization efficiency—especially in scenarios where mature PMUs are unavailable.
📝 Abstract
As RISC-V architectures proliferate across embedded and high-performance domains, developers face persistent challenges in performance optimization due to fragmented tooling, immature hardware features, and platform-specific defects. This paper delivers a pragmatic methodology for extracting actionable performance insights on RISC-V systems, even under constrained or unreliable hardware conditions. We present a workaround to circumvent hardware bugs in one of the popular RISC-V implementations, enabling robust event sampling. For memory-compute bottleneck analysis, we introduce compiler-driven Roofline tooling that operates without hardware PMU dependencies, leveraging LLVM-based instrumentation to derive operational intensity and throughput metrics directly from application IR. Our open source toolchain automates these workarounds, unifying PMU data correction and compiler-guided Roofline construction into a single workflow.