accelerated computing

Designs, implements, and evaluates software, algorithms, and system components that use hardware accelerators (e.g., GPUs, TPUs, FPGAs) to increase computational throughput and efficiency; this includes writing accelerator kernels, managing memory and data movement, coordinating CPU–accelerator orchestration, and building runtime and scheduling strategies. It also encompasses measuring and optimizing performance, scalability, resource utilization, and numerical/precision trade-offs for accelerated workloads.

acceleratedcomputing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.76
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

FPGA or GPU? Analyzing Comparative Research for Application-Specific Guidance

Mar 22, 2025
AA
Arnab A Purkayastha
🏛️ Western New England University | The Citadel

Existing FPGA–GPU comparative studies predominantly focus on raw performance metrics and lack domain-specific guidance for accelerator selection. Method: This paper proposes an application-oriented, fine-grained comparative framework that systematically synthesizes over 100 studies, conducting cross-domain (e.g., AI, HPC, network processing, scientific computing) classification and cross-evaluation along three dimensions: performance, energy efficiency, and programmability. Contribution/Results: The study innovatively establishes the first empirically grounded applicability boundaries for FPGAs and GPUs: FPGAs excel in low-latency, high-throughput customized pipelines and energy-constrained scenarios; GPUs are superior for massively parallel, computation-intensive workloads with stable algorithms. The resulting actionable decision-making guide enables researchers and engineers to select hardware accelerators based on domain-specific requirements, thereby bridging the gap between architectural characteristics and real-world application needs.

Analyzing FPGA vs GPU performance for application-specific guidanceBridging research gap on ideal accelerators for domain-specific applicationsProviding actionable recommendations for hardware accelerator selection

Toward Heterogeneous, Distributed, and Energy-Efficient Computing with SYCL

May 09, 2025
BC
Biagio Cosenza
🏛️ University of Salerno | Queen's University Belfast | University of Innsbruck

To address the challenges of programming complexity for GPU/FPGA accelerators, high cross-node data movement overhead, and difficulty in energy-efficiency optimization in heterogeneous distributed HPC systems, this paper proposes a high-level programming framework based on SYCL 2020. Its core contributions are: (1) Celerity—a novel distributed task dispatching mechanism supporting standard SYCL semantics, enabling unified scheduling and load balancing across CPUs, GPUs, and FPGAs; and (2) SYnergy—a power-modeling-driven co-optimization extension that integrates feedback control with multi-level memory-aware task mapping to achieve energy-aware execution. The framework is fully compatible with mainstream SYCL implementations and requires no modifications to existing SYCL code. Experimental evaluation on heterogeneous clusters demonstrates up to a 2.3× improvement in energy efficiency and a 1.8× speedup in task dispatching throughput.

Achieving energy-efficient computing in heterogeneous systemsEfficient programming of GPUs and accelerators in HPCHandling data movement between distributed computing nodes

Evaluating Rapid Makespan Predictions for Heterogeneous Systems with Programmable Logic

Oct 08, 2025
MW
Martin Wilhelm
🏛️ Otto-von-Guericke University | University of Applied Sciences

To address the challenge of rapidly evaluating the impact of dynamic task mapping adjustments on overall makespan in heterogeneous systems (CPU/GPU/FPGA), this paper proposes a lightweight prediction framework integrating abstract task graph modeling, empirical performance profiling, and analytical function fitting. The method explicitly models critical high-level factors—including inter-device data transfer overhead and hardware resource congestion—thereby bridging theoretical analysis and measured performance. Compared to conventional analytical models, our framework significantly improves cross-platform makespan prediction accuracy and is systematically validated on real heterogeneous hardware. Key contributions are: (1) the first unified execution time prediction framework supporting multiple hardware backends; (2) empirical identification of data transfer and resource congestion as dominant sources of prediction error; and (3) a scalable, empirically grounded foundation for rapid, reliable mapping decisions.

Bridging analytical predictions with real-world heterogeneous system performanceEvaluating data transfer and device congestion challenges in acceleratorsPredicting makespan impact of task mapping changes in heterogeneous systems

This work addresses the inefficiency and suboptimal energy consumption of data aggregation operations on heterogeneous hardware platforms. To bridge this gap, the authors propose a hybrid hardware acceleration framework that synergistically combines unified abstractions with platform-specific optimizations, effectively balancing programmability, portability, and architectural specialization across CPUs, GPUs, and FPGAs. By introducing a common abstraction layer while incorporating tailored optimization strategies for each hardware target, the approach achieves significant improvements in both performance and energy efficiency across all three mainstream architectures. The evaluation demonstrates consistent gains not only in device-level computation but also in end-to-end processing metrics, thereby establishing an effective trade-off between generality and high performance for data aggregation workloads.

aggregationdata analyticsFPGA

This work addresses the inefficiency of existing programmable architectures in handling sparse or irregular data and the inflexibility of dedicated accelerators when confronted with new kernels or input patterns. To bridge this gap, the paper proposes Canon, a novel architecture that integrates a programmable finite state machine (FSM) with a dynamic, data-driven execution orchestration mechanism to generate control flow at runtime. Canon further introduces a time-interleaved SIMD execution model that constructs an evolving dataflow to maximize parallelism. This design achieves performance and energy efficiency approaching that of specialized accelerators across a range of data-oblivious and data-driven kernels, while preserving the programmability and flexibility of general-purpose architectures.

execution orchestrationirregular dataperformance fragility

Latest Papers

What's happening recently
View more

Enabling Heterogeneous Performance Analysis for Scientific Workloads

Nov 17, 2025
MG
Maksymilian Graczyk
🏛️ CERN | HES-SO Valais-Wallis

Addressing the challenge of jointly optimizing performance and energy efficiency for scientific workloads on heterogeneous systems (CPU/GPU/FPGA), this paper introduces Adaptyst—an open-source, architecture-agnostic performance analysis framework. Methodologically, it pioneers the deep integration of eBPF with uprobes (dynamic instrumentation) and USDT (user-space static tracing), enabling cross-architecture, low-overhead, high-fidelity fine-grained runtime behavior monitoring and performance data collection. Through systematic evaluation of the overhead, accuracy, and integrability of both eBPF probe mechanisms, the study delineates their applicability boundaries and optimization strategies in heterogeneous environments. Experiments demonstrate that Adaptyst effectively supports intelligent task-to-accelerator scheduling decisions by identifying optimal compute units, thereby establishing a novel paradigm for heterogeneous performance analysis and delivering a reusable, production-ready infrastructure.

Developing architecture-agnostic analysis methods for scientific workloadsEnabling performance analysis for heterogeneous computing systemsEvaluating eBPF-based methods for future integration into Adaptyst

This work proposes the first modular profiling framework tailored for hardware accelerators, addressing the lack of low-overhead and flexible program analysis tools in modern computing systems. By abstracting underlying performance APIs and integrating with mainstream deep learning frameworks, the framework offers a unified interface to capture runtime events across multiple abstraction levels and enables rapid prototyping. It features a GPU-accelerated backend, multi-level event tracing, and cross-platform compatibility (NVIDIA/AMD), achieving high scalability and minimal profiling overhead in both single- and multi-GPU settings. Experimental results demonstrate that, on representative deep learning workloads, the framework achieves up to 1.3×10⁴ times faster profiling compared to conventional tools while delivering fine-grained performance insights.

hardware acceleratorslow-overheadmodular framework

This work addresses the lack of automated, efficient methods for evaluating how closely existing benchmarks resemble real-world high-performance computing (HPC) applications in terms of hardware performance characteristics. The authors propose a novel performance similarity metric based on hardware usage patterns, introducing for the first time two distinct classes of computational kernels that exhibit similar performance behavior. They develop a scalable, automated evaluation framework that integrates performance feature analysis, kernel classification, and cross-platform (CPU/GPU) similarity assessment. The effectiveness and practicality of this approach are demonstrated by accurately matching computational kernels from the Kripke proxy application to those in the RAJA Performance Suite, thereby validating the method’s capability to identify functionally analogous kernels across diverse hardware architectures.

code representationcomputational kernelsHPC benchmarks

Traditional efficiency metrics struggle to accurately assess resource utilization in heterogeneous high-performance computing systems that combine CPUs and accelerators. This work extends the POP efficiency model by introducing a hardware-agnostic, host-device dual-branch hierarchical efficiency framework. It uniquely defines a multiplicative efficiency decomposition on the device side, symmetric to that on the host, separately capturing mixed execution/offload efficiency and device parallel efficiency. Implemented via the lightweight TALP monitoring library, the approach supports both runtime and post-mortem analysis and outputs results in human-readable and machine-readable formats. Experiments on synthetic benchmarks and three real-world HPC applications demonstrate that the proposed methodology effectively uncovers performance bottlenecks related to offloading, load balancing, and task scheduling, offering developers actionable insights for optimization.

acceleratorsefficiency analysisheterogeneous computing

Deploying GEMM on tile-based multi-PE accelerators faces challenges of deployment complexity and deep hardware-software coupling. To address this, we propose an end-to-end automated deployment framework. Our approach introduces the novel “Design in Tiles” paradigm, integrating configurable execution modeling, hardware-aware automatic mapping, hierarchical tiling scheduling, and compute-memory co-optimized compilation. For the first time, we achieve superior PE utilization over NVIDIA GH200’s expert-tuned library on a large-scale 32×32 tile configuration. At FP8 precision, our framework delivers 1979 TFLOPS peak performance and accelerates diverse matrix shapes by 1.2–2.0× relative to GH200. This work bridges the compilation gap between configurable hardware architectures and high-level computational graphs, establishing a general, efficient, and scalable methodology for automatic mapping onto domain-specific accelerators.

Addresses programming difficulty due to hardware-software couplingAutomates GEMM deployment on tile-based many-PE acceleratorsImproves performance over expert-tuned libraries on large configurations