Score
Designs, implements, and evaluates software, algorithms, and system components that use hardware accelerators (e.g., GPUs, TPUs, FPGAs) to increase computational throughput and efficiency; this includes writing accelerator kernels, managing memory and data movement, coordinating CPU–accelerator orchestration, and building runtime and scheduling strategies. It also encompasses measuring and optimizing performance, scalability, resource utilization, and numerical/precision trade-offs for accelerated workloads.
Existing FPGA–GPU comparative studies predominantly focus on raw performance metrics and lack domain-specific guidance for accelerator selection. Method: This paper proposes an application-oriented, fine-grained comparative framework that systematically synthesizes over 100 studies, conducting cross-domain (e.g., AI, HPC, network processing, scientific computing) classification and cross-evaluation along three dimensions: performance, energy efficiency, and programmability. Contribution/Results: The study innovatively establishes the first empirically grounded applicability boundaries for FPGAs and GPUs: FPGAs excel in low-latency, high-throughput customized pipelines and energy-constrained scenarios; GPUs are superior for massively parallel, computation-intensive workloads with stable algorithms. The resulting actionable decision-making guide enables researchers and engineers to select hardware accelerators based on domain-specific requirements, thereby bridging the gap between architectural characteristics and real-world application needs.
To address the challenges of programming complexity for GPU/FPGA accelerators, high cross-node data movement overhead, and difficulty in energy-efficiency optimization in heterogeneous distributed HPC systems, this paper proposes a high-level programming framework based on SYCL 2020. Its core contributions are: (1) Celerity—a novel distributed task dispatching mechanism supporting standard SYCL semantics, enabling unified scheduling and load balancing across CPUs, GPUs, and FPGAs; and (2) SYnergy—a power-modeling-driven co-optimization extension that integrates feedback control with multi-level memory-aware task mapping to achieve energy-aware execution. The framework is fully compatible with mainstream SYCL implementations and requires no modifications to existing SYCL code. Experimental evaluation on heterogeneous clusters demonstrates up to a 2.3× improvement in energy efficiency and a 1.8× speedup in task dispatching throughput.
To address the challenge of rapidly evaluating the impact of dynamic task mapping adjustments on overall makespan in heterogeneous systems (CPU/GPU/FPGA), this paper proposes a lightweight prediction framework integrating abstract task graph modeling, empirical performance profiling, and analytical function fitting. The method explicitly models critical high-level factors—including inter-device data transfer overhead and hardware resource congestion—thereby bridging theoretical analysis and measured performance. Compared to conventional analytical models, our framework significantly improves cross-platform makespan prediction accuracy and is systematically validated on real heterogeneous hardware. Key contributions are: (1) the first unified execution time prediction framework supporting multiple hardware backends; (2) empirical identification of data transfer and resource congestion as dominant sources of prediction error; and (3) a scalable, empirically grounded foundation for rapid, reliable mapping decisions.
This work addresses the inefficiency and suboptimal energy consumption of data aggregation operations on heterogeneous hardware platforms. To bridge this gap, the authors propose a hybrid hardware acceleration framework that synergistically combines unified abstractions with platform-specific optimizations, effectively balancing programmability, portability, and architectural specialization across CPUs, GPUs, and FPGAs. By introducing a common abstraction layer while incorporating tailored optimization strategies for each hardware target, the approach achieves significant improvements in both performance and energy efficiency across all three mainstream architectures. The evaluation demonstrates consistent gains not only in device-level computation but also in end-to-end processing metrics, thereby establishing an effective trade-off between generality and high performance for data aggregation workloads.
This work addresses the inefficiency of existing programmable architectures in handling sparse or irregular data and the inflexibility of dedicated accelerators when confronted with new kernels or input patterns. To bridge this gap, the paper proposes Canon, a novel architecture that integrates a programmable finite state machine (FSM) with a dynamic, data-driven execution orchestration mechanism to generate control flow at runtime. Canon further introduces a time-interleaved SIMD execution model that constructs an evolving dataflow to maximize parallelism. This design achieves performance and energy efficiency approaching that of specialized accelerators across a range of data-oblivious and data-driven kernels, while preserving the programmability and flexibility of general-purpose architectures.
Addressing the challenge of jointly optimizing performance and energy efficiency for scientific workloads on heterogeneous systems (CPU/GPU/FPGA), this paper introduces Adaptyst—an open-source, architecture-agnostic performance analysis framework. Methodologically, it pioneers the deep integration of eBPF with uprobes (dynamic instrumentation) and USDT (user-space static tracing), enabling cross-architecture, low-overhead, high-fidelity fine-grained runtime behavior monitoring and performance data collection. Through systematic evaluation of the overhead, accuracy, and integrability of both eBPF probe mechanisms, the study delineates their applicability boundaries and optimization strategies in heterogeneous environments. Experiments demonstrate that Adaptyst effectively supports intelligent task-to-accelerator scheduling decisions by identifying optimal compute units, thereby establishing a novel paradigm for heterogeneous performance analysis and delivering a reusable, production-ready infrastructure.
This work proposes the first modular profiling framework tailored for hardware accelerators, addressing the lack of low-overhead and flexible program analysis tools in modern computing systems. By abstracting underlying performance APIs and integrating with mainstream deep learning frameworks, the framework offers a unified interface to capture runtime events across multiple abstraction levels and enables rapid prototyping. It features a GPU-accelerated backend, multi-level event tracing, and cross-platform compatibility (NVIDIA/AMD), achieving high scalability and minimal profiling overhead in both single- and multi-GPU settings. Experimental results demonstrate that, on representative deep learning workloads, the framework achieves up to 1.3×10⁴ times faster profiling compared to conventional tools while delivering fine-grained performance insights.
This work addresses the lack of automated, efficient methods for evaluating how closely existing benchmarks resemble real-world high-performance computing (HPC) applications in terms of hardware performance characteristics. The authors propose a novel performance similarity metric based on hardware usage patterns, introducing for the first time two distinct classes of computational kernels that exhibit similar performance behavior. They develop a scalable, automated evaluation framework that integrates performance feature analysis, kernel classification, and cross-platform (CPU/GPU) similarity assessment. The effectiveness and practicality of this approach are demonstrated by accurately matching computational kernels from the Kripke proxy application to those in the RAJA Performance Suite, thereby validating the method’s capability to identify functionally analogous kernels across diverse hardware architectures.
Traditional efficiency metrics struggle to accurately assess resource utilization in heterogeneous high-performance computing systems that combine CPUs and accelerators. This work extends the POP efficiency model by introducing a hardware-agnostic, host-device dual-branch hierarchical efficiency framework. It uniquely defines a multiplicative efficiency decomposition on the device side, symmetric to that on the host, separately capturing mixed execution/offload efficiency and device parallel efficiency. Implemented via the lightweight TALP monitoring library, the approach supports both runtime and post-mortem analysis and outputs results in human-readable and machine-readable formats. Experiments on synthetic benchmarks and three real-world HPC applications demonstrate that the proposed methodology effectively uncovers performance bottlenecks related to offloading, load balancing, and task scheduling, offering developers actionable insights for optimization.
Deploying GEMM on tile-based multi-PE accelerators faces challenges of deployment complexity and deep hardware-software coupling. To address this, we propose an end-to-end automated deployment framework. Our approach introduces the novel “Design in Tiles” paradigm, integrating configurable execution modeling, hardware-aware automatic mapping, hierarchical tiling scheduling, and compute-memory co-optimized compilation. For the first time, we achieve superior PE utilization over NVIDIA GH200’s expert-tuned library on a large-scale 32×32 tile configuration. At FP8 precision, our framework delivers 1979 TFLOPS peak performance and accelerates diverse matrix shapes by 1.2–2.0× relative to GH200. This work bridges the compilation gap between configurable hardware architectures and high-level computational graphs, establishing a general, efficient, and scalable methodology for automatic mapping onto domain-specific accelerators.