Score
Designs, implements, and analyzes system architectures that integrate hardware accelerators—such as DPUs, TPUs, BlueField devices, and other offload engines—into servers and clusters, including DPU/TPU offload mechanisms, transport-level offloads, and multi-level offloading strategies. Builds orchestration and cluster designs, implements hardware offload interfaces, and carries out performance profiling and workload optimization to evaluate and tune end-to-end accelerator-enabled deployments.
This paper addresses key challenges in building heterogeneous computing systems through DPU/SmartNIC–CPU co-design. We conduct a systematic survey of over 100 representative works published between 2018 and 2024. Methodologically, we propose the first comprehensive taxonomy for DPU–CPU collaborative computing, categorizing research along three dimensions: hardware architectures (e.g., NVIDIA BlueField, Pensando), programming models (e.g., eBPF, DPDK, SPDK), and offloading mechanisms with coordinated scheduling techniques. Our analysis identifies driving forces behind technological evolution, fundamental bottlenecks—including memory consistency, inter-device communication latency, and software stack fragmentation—and emerging trends toward tighter hardware–software integration. As a result, we construct a domain knowledge graph spanning architectural principles, software stacks, and application scenarios (e.g., AI/ML acceleration, cloud data centers). This work establishes an authoritative benchmark and methodological foundation for co-designed DPU hardware/software development, performance modeling, and domain-specific adaptation.
SmartNIC Data Processing Units (DPUs) offer a promising solution for saving high-end CPU resources by offloading tasks to programmable cores near the network interface. In this work, we explore the feasibility of SmartNIC DPUs in supporting an asynchronous communication model called "fire-and-forget", particularly its core message routing service. We design a communication offloading engine called Buddy that decouples communication tasks from the application process. Buddy runs flexibly on SmartNIC DPUs such as the Nvidia BlueField-3 DPU and generic x86 CPUs. Our evaluation results in five applications identify the memory-to-communication ratio as a key predictor of the offloading performance. Host-dominated workloads, such as Quicksilver and Sparse Matrix Transpose, achieved up to 1.55x speedup with communication offloaded to the DPU. We further identify a 625x increase in DRAM traffic due to the absence of Direct Cache Access support on the DPU, highlighting a critical need in future SmartNIC designs.
Existing benchmarks lack systematic evaluation of Data Processing Unit (DPU) capabilities for data-intensive workloads. Method: We propose the first DPU-specific benchmark suite tailored for data processing, built upon a scalable abstraction framework that uniquely supports heterogeneous DPU architectures and multiple data processing stacks—including network I/O, memory bandwidth, coprocessor acceleration, and storage offloading—enabling cross-platform, modular performance assessment. Contribution/Results: Evaluated across mainstream DPU platforms, our suite demonstrates 1.8×–5.3× throughput improvements over CPUs on representative workloads such as query processing, compression, and encryption. It is the first to quantitatively characterize the performance benefits and fundamental bottlenecks of DPU offloading, thereby establishing a rigorous foundation for DPU data-processing evaluation and closing a critical gap in the systems benchmarking landscape.
Hardware disaggregation aims to transcend traditional server boundaries and establish a unified resource pool spanning cabinets or racks, yet faces critical challenges in resource pooling and coordinated scheduling, energy-efficiency optimization, and system-level trade-offs. This paper proposes a cross-layer co-optimization framework integrating system architecture design, resource pooling mechanisms, fine-grained scheduling algorithms, and a multi-objective energy-efficiency evaluation model. It systematically reveals the deep impacts of decoupled architectures on application development, hardware configuration, and power/thermal management. Through numerical modeling and quantitative analysis, we first characterize the three-dimensional trade-off among pooling granularity, scheduling overhead, and energy efficiency—filling a key gap in pooling-scheduling co-optimization research. Experiments demonstrate that our architecture improves resource utilization by 32–47%, reduces Power Usage Effectiveness (PUE) by 0.08–0.15, and significantly enhances adaptability to heterogeneous workloads.
To address high-latency accelerator access caused by host involvement, this paper proposes OffRAC—the first host-free remote accelerator direct-call paradigm—elevating hardware accelerators (e.g., FPGAs) to programmable, native first-class computing resources in the network. Methodologically, OffRAC integrates serverless stateless function abstraction with hardware-level multi-tenancy isolation, enabling a lightweight remote invocation protocol, request aggregation mechanism, and dedicated scheduling framework. Its key innovation lies in eliminating the traditional host-mediated bottleneck, thereby achieving low-overhead, strongly isolated, cross-client direct function invocation. Evaluated on a real FPGA platform, the prototype achieves an end-to-end invocation latency of 10.5 μs and throughput of 85 Gbps. OffRAC significantly improves scalability, energy efficiency, and resource utilization in ultra-low-latency data processing scenarios.
Addressing the challenge of jointly optimizing performance and energy efficiency for scientific workloads on heterogeneous systems (CPU/GPU/FPGA), this paper introduces Adaptyst—an open-source, architecture-agnostic performance analysis framework. Methodologically, it pioneers the deep integration of eBPF with uprobes (dynamic instrumentation) and USDT (user-space static tracing), enabling cross-architecture, low-overhead, high-fidelity fine-grained runtime behavior monitoring and performance data collection. Through systematic evaluation of the overhead, accuracy, and integrability of both eBPF probe mechanisms, the study delineates their applicability boundaries and optimization strategies in heterogeneous environments. Experiments demonstrate that Adaptyst effectively supports intelligent task-to-accelerator scheduling decisions by identifying optimal compute units, thereby establishing a novel paradigm for heterogeneous performance analysis and delivering a reusable, production-ready infrastructure.
The deployment of high-density AI accelerators renders traditional data center power delivery hierarchies inefficient in utilizing provisioned power, leading to resource waste and stranded capacity. This work presents the first integrated evaluation framework that jointly models GPU, compute, and storage placement, leveraging real-world Azure traces of workload arrivals, overbooking patterns, and hardware retirement schedules to co-optimize power, performance, and cost. Innovatively incorporating multi-resource stranding into power infrastructure design, the study proposes “deployable capacity”—the actual computational capacity that can be effectively powered and utilized—as a more meaningful planning objective than conventional “nameplate power.” It further quantifies the impact of high-density AI systems on deployable capacity, effective capital expenditure, and delivered performance, offering critical insights for rethinking data center power architectures in the AI era.
This work addresses the inefficiency and suboptimal energy consumption of data aggregation operations on heterogeneous hardware platforms. To bridge this gap, the authors propose a hybrid hardware acceleration framework that synergistically combines unified abstractions with platform-specific optimizations, effectively balancing programmability, portability, and architectural specialization across CPUs, GPUs, and FPGAs. By introducing a common abstraction layer while incorporating tailored optimization strategies for each hardware target, the approach achieves significant improvements in both performance and energy efficiency across all three mainstream architectures. The evaluation demonstrates consistent gains not only in device-level computation but also in end-to-end processing metrics, thereby establishing an effective trade-off between generality and high performance for data aggregation workloads.
This work addresses the challenge of efficiently exploring the vast and physically constrained design space of cross-layer heterogeneous systems to support mixed AI and high-performance computing (HPC) workloads. To this end, the authors propose CHASE, a novel framework that decouples hardware architecture design from task mapping. CHASE leverages hierarchical type graphs for system modeling, a topology-aware mapper, and a telemetry-guided optimizer to enable application-driven architecture search under deployment constraints. Experimental results demonstrate that CHASE achieves geometric mean speedups of 6.20× and 2.12× on sparse computing and large language model workloads, respectively, while reducing mapping time by 60.5% on average and converging to near-global-optimal solutions within 64 iterations.
This work addresses performance bottlenecks in existing servers during microsecond- to millisecond-scale hardware offloading, which often stem from context-switching overhead or busy-waiting. The authors propose a fine-grained offloading approach that requires no modification to server code, leveraging for the first time the server’s built-in suspend-resume concurrency mechanism and reframing offload scheduling as a routing problem. By injecting a fiber runtime via LD_PRELOAD, integrating native deferred response handling with executor submission, and employing page-protection techniques to ensure atomicity and safety, the method achieves significant speedups with only 22–138 lines of adapter code. Evaluated across ten widely used servers, it delivers 1.2–5.4× acceleration; notably, it attains a 17.3× speedup on unmodified thread-per-connection binaries and demonstrates both safety and efficacy in Redis.