hardware accelerator integration

Designs, implements, and analyzes system architectures that integrate hardware accelerators—such as DPUs, TPUs, BlueField devices, and other offload engines—into servers and clusters, including DPU/TPU offload mechanisms, transport-level offloads, and multi-level offloading strategies. Builds orchestration and cluster designs, implements hardware offload interfaces, and carries out performance profiling and workload optimization to evaluate and tune end-to-end accelerator-enabled deployments.

hardwareacceleratorintegration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.69
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$228K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

SmartNIC Data Processing Units (DPUs) offer a promising solution for saving high-end CPU resources by offloading tasks to programmable cores near the network interface. In this work, we explore the feasibility of SmartNIC DPUs in supporting an asynchronous communication model called "fire-and-forget", particularly its core message routing service. We design a communication offloading engine called Buddy that decouples communication tasks from the application process. Buddy runs flexibly on SmartNIC DPUs such as the Nvidia BlueField-3 DPU and generic x86 CPUs. Our evaluation results in five applications identify the memory-to-communication ratio as a key predictor of the offloading performance. Host-dominated workloads, such as Quicksilver and Sparse Matrix Transpose, achieved up to 1.55x speedup with communication offloaded to the DPU. We further identify a 625x increase in DRAM traffic due to the absence of Direct Cache Access support on the DPU, highlighting a critical need in future SmartNIC designs.

Communication OffloadingCPU Resource SavingFire-and-Forget

dpBento: Benchmarking DPUs for Data Processing

Apr 07, 2025
JH
Jiasheng Hu
🏛️ University of Toronto | Microsoft Research | National University of Singapore

Existing benchmarks lack systematic evaluation of Data Processing Unit (DPU) capabilities for data-intensive workloads. Method: We propose the first DPU-specific benchmark suite tailored for data processing, built upon a scalable abstraction framework that uniquely supports heterogeneous DPU architectures and multiple data processing stacks—including network I/O, memory bandwidth, coprocessor acceleration, and storage offloading—enabling cross-platform, modular performance assessment. Contribution/Results: Evaluated across mainstream DPU platforms, our suite demonstrates 1.8×–5.3× throughput improvements over CPUs on representative workloads such as query processing, compression, and encryption. It is the first to quantitatively characterize the performance benefits and fundamental bottlenecks of DPU offloading, thereby establishing a rigorous foundation for DPU data-processing evaluation and closing a critical gap in the systems benchmarking landscape.

Assessing performance impacts of offloading to DPUsBenchmarking DPUs for diverse data processing tasksLack of comprehensive DPU benchmarking tools

Disaggregated Architectures and the Redesign of Data Center Ecosystems: Scheduling, Pooling, and Infrastructure Trade-offs

Nov 06, 2025
CG
Chao Guo
🏛️ Centre for Intelligent Multidimensional Data Analysis Limited | City University of Hong Kong

Hardware disaggregation aims to transcend traditional server boundaries and establish a unified resource pool spanning cabinets or racks, yet faces critical challenges in resource pooling and coordinated scheduling, energy-efficiency optimization, and system-level trade-offs. This paper proposes a cross-layer co-optimization framework integrating system architecture design, resource pooling mechanisms, fine-grained scheduling algorithms, and a multi-objective energy-efficiency evaluation model. It systematically reveals the deep impacts of decoupled architectures on application development, hardware configuration, and power/thermal management. Through numerical modeling and quantitative analysis, we first characterize the three-dimensional trade-off among pooling granularity, scheduling overhead, and energy efficiency—filling a key gap in pooling-scheduling co-optimization research. Experiments demonstrate that our architecture improves resource utilization by 32–47%, reduces Power Usage Effectiveness (PUE) by 0.08–0.15, and significantly enhances adaptability to heterogeneous workloads.

Addressing scheduling and pooling challenges in data centersOptimizing hardware configuration and power systemsTransforming server fleets into unified resource pools

OffRAC: Offloading Through Remote Accelerator Calls

Apr 06, 2025
ZY
Ziyi Yang
🏛️ KAUST | Microsoft Reserch | Technical University of Darmstadt

To address high-latency accelerator access caused by host involvement, this paper proposes OffRAC—the first host-free remote accelerator direct-call paradigm—elevating hardware accelerators (e.g., FPGAs) to programmable, native first-class computing resources in the network. Methodologically, OffRAC integrates serverless stateless function abstraction with hardware-level multi-tenancy isolation, enabling a lightweight remote invocation protocol, request aggregation mechanism, and dedicated scheduling framework. Its key innovation lies in eliminating the traditional host-mediated bottleneck, thereby achieving low-overhead, strongly isolated, cross-client direct function invocation. Evaluated on a real FPGA platform, the prototype achieves an end-to-end invocation latency of 10.5 μs and throughput of 85 Gbps. OffRAC significantly improves scalability, energy efficiency, and resource utilization in ultra-low-latency data processing scenarios.

Enabling direct remote calls to FPGA-based accelerators efficientlyEnsuring performance isolation and scalability for multi-tenant accelerator useReducing latency in accelerator access by eliminating host involvement

Latest Papers

What's happening recently
View more

Enabling Heterogeneous Performance Analysis for Scientific Workloads

Nov 17, 2025
MG
Maksymilian Graczyk
🏛️ CERN | HES-SO Valais-Wallis

Addressing the challenge of jointly optimizing performance and energy efficiency for scientific workloads on heterogeneous systems (CPU/GPU/FPGA), this paper introduces Adaptyst—an open-source, architecture-agnostic performance analysis framework. Methodologically, it pioneers the deep integration of eBPF with uprobes (dynamic instrumentation) and USDT (user-space static tracing), enabling cross-architecture, low-overhead, high-fidelity fine-grained runtime behavior monitoring and performance data collection. Through systematic evaluation of the overhead, accuracy, and integrability of both eBPF probe mechanisms, the study delineates their applicability boundaries and optimization strategies in heterogeneous environments. Experiments demonstrate that Adaptyst effectively supports intelligent task-to-accelerator scheduling decisions by identifying optimal compute units, thereby establishing a novel paradigm for heterogeneous performance analysis and delivering a reusable, production-ready infrastructure.

Developing architecture-agnostic analysis methods for scientific workloadsEnabling performance analysis for heterogeneous computing systemsEvaluating eBPF-based methods for future integration into Adaptyst

The deployment of high-density AI accelerators renders traditional data center power delivery hierarchies inefficient in utilizing provisioned power, leading to resource waste and stranded capacity. This work presents the first integrated evaluation framework that jointly models GPU, compute, and storage placement, leveraging real-world Azure traces of workload arrivals, overbooking patterns, and hardware retirement schedules to co-optimize power, performance, and cost. Innovatively incorporating multi-resource stranding into power infrastructure design, the study proposes “deployable capacity”—the actual computational capacity that can be effectively powered and utilized—as a more meaningful planning objective than conventional “nameplate power.” It further quantifies the impact of high-density AI systems on deployable capacity, effective capital expenditure, and delivered performance, offering critical insights for rethinking data center power architectures in the AI era.

AI acceleratorsdatacenterpower delivery

This work addresses the inefficiency and suboptimal energy consumption of data aggregation operations on heterogeneous hardware platforms. To bridge this gap, the authors propose a hybrid hardware acceleration framework that synergistically combines unified abstractions with platform-specific optimizations, effectively balancing programmability, portability, and architectural specialization across CPUs, GPUs, and FPGAs. By introducing a common abstraction layer while incorporating tailored optimization strategies for each hardware target, the approach achieves significant improvements in both performance and energy efficiency across all three mainstream architectures. The evaluation demonstrates consistent gains not only in device-level computation but also in end-to-end processing metrics, thereby establishing an effective trade-off between generality and high performance for data aggregation workloads.

aggregationdata analyticsFPGA

This work addresses the challenge of efficiently exploring the vast and physically constrained design space of cross-layer heterogeneous systems to support mixed AI and high-performance computing (HPC) workloads. To this end, the authors propose CHASE, a novel framework that decouples hardware architecture design from task mapping. CHASE leverages hierarchical type graphs for system modeling, a topology-aware mapper, and a telemetry-guided optimizer to enable application-driven architecture search under deployment constraints. Experimental results demonstrate that CHASE achieves geometric mean speedups of 6.20× and 2.12× on sparse computing and large language model workloads, respectively, while reducing mapping time by 60.5% on average and converging to near-global-optimal solutions within 64 iterations.

Architecture ExplorationCross-layer Heterogeneous SystemDeployment Constraints

This work addresses performance bottlenecks in existing servers during microsecond- to millisecond-scale hardware offloading, which often stem from context-switching overhead or busy-waiting. The authors propose a fine-grained offloading approach that requires no modification to server code, leveraging for the first time the server’s built-in suspend-resume concurrency mechanism and reframing offload scheduling as a routing problem. By injecting a fiber runtime via LD_PRELOAD, integrating native deferred response handling with executor submission, and employing page-protection techniques to ensure atomicity and safety, the method achieves significant speedups with only 22–138 lines of adapter code. Evaluated across ten widely used servers, it delivers 1.2–5.4× acceleration; notably, it attains a 17.3× speedup on unmodified thread-per-connection binaries and demonstrates both safety and efficacy in Redis.

computation offloadconcurrencyfine-grained

Hot Scholars

OB

Oliver Bringmann

Professor of Embedded Systems, Eberhard Karls Universität Tübingen, Germany
Embedded System DesignSystem Modeling and SimulationAutomotive ElectronicsSafety-critical Systems
LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
JZ

Jingyang Zhu

Ph.D. Student in Shanghaitech University
Edge AINetworkingSatellite
KB

Khaled B. Letaief

Member of US National Academy of Engineering and New Bright Professor of Engineering, HKUST
Wirelesscommunications
YS

Yuanming Shi

Professor, ShanghaiTech University
Space Computing NetworksEdge Artificial IntelligenceLarge-Scale Optimization