π€ AI Summary
Addressing the challenge of jointly optimizing performance and energy efficiency for scientific workloads on heterogeneous systems (CPU/GPU/FPGA), this paper introduces Adaptystβan open-source, architecture-agnostic performance analysis framework. Methodologically, it pioneers the deep integration of eBPF with uprobes (dynamic instrumentation) and USDT (user-space static tracing), enabling cross-architecture, low-overhead, high-fidelity fine-grained runtime behavior monitoring and performance data collection. Through systematic evaluation of the overhead, accuracy, and integrability of both eBPF probe mechanisms, the study delineates their applicability boundaries and optimization strategies in heterogeneous environments. Experiments demonstrate that Adaptyst effectively supports intelligent task-to-accelerator scheduling decisions by identifying optimal compute units, thereby establishing a novel paradigm for heterogeneous performance analysis and delivering a reusable, production-ready infrastructure.
π Abstract
Heterogeneous computing integrates diverse processing elements, such as CPUs, GPUs, and FPGAs, within a single system, aiming to leverage the strengths of each architecture to optimize performance and energy consumption. In this context, efficient performance analysis plays a critical role in determining the most suitable platform for dispatching tasks, ensuring that workloads are allocated to the processing units where they can execute most effectively. Adaptyst is a novel ongoing effort at CERN, with the aim to develop an open-source, architecture-agnostic performance analysis for scientific workloads. This study explores the performance and implementation complexity of two built-in eBPF-based methods such as Uprobes and USDT, with the aim of outlining a roadmap for future integration into Adaptyst and advancing toward heterogeneous performance analysis capabilities.