OffRAC: Offloading Through Remote Accelerator Calls

📅 2025-04-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address high-latency accelerator access caused by host involvement, this paper proposes OffRAC—the first host-free remote accelerator direct-call paradigm—elevating hardware accelerators (e.g., FPGAs) to programmable, native first-class computing resources in the network. Methodologically, OffRAC integrates serverless stateless function abstraction with hardware-level multi-tenancy isolation, enabling a lightweight remote invocation protocol, request aggregation mechanism, and dedicated scheduling framework. Its key innovation lies in eliminating the traditional host-mediated bottleneck, thereby achieving low-overhead, strongly isolated, cross-client direct function invocation. Evaluated on a real FPGA platform, the prototype achieves an end-to-end invocation latency of 10.5 μs and throughput of 85 Gbps. OffRAC significantly improves scalability, energy efficiency, and resource utilization in ultra-low-latency data processing scenarios.

Technology Category

Machine Learning: Hardware-aware MLSearch and Optimization: Distributed SearchPlanning, Routing, and Scheduling: Optimization of Spatio-temporal Systems

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Experiences and lessons learnt from Web-based algorithms and system deploymentsSearch and Retrieval-Augmented AI: Personalized, context-aware and across-device searchUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendation
📝 Abstract
Modern applications increasingly demand ultra-low latency for data processing, often facilitated by host-controlled accelerators like GPUs and FPGAs. However, significant delays result from host involvement in accessing accelerators. To address this limitation, we introduce a novel paradigm we call Offloading through Remote Accelerator Calls (OffRAC), which elevates accelerators to first-class compute resources. OffRAC enables direct calls to FPGA-based accelerators without host involvement. Utilizing the stateless function abstraction of serverless computing, with applications decomposed into simpler stateless functions, offloading promotes efficient acceleration and distribution of computational loads across the network. To realize this proposal, we present a prototype design and implementation of an OffRAC platform for FPGAs that assembles diverse requests from multiple clients into complete accelerator calls with multi-tenancy performance isolation. This design minimizes the implementation complexity for accelerator users while ensuring isolation and programmability. Results show that the OffRAC approach reduces the latency of network calls to accelerators down to approximately 10.5 us, as well as sustaining high application throughput up to 85Gbps, demonstrating scalability and efficiency, making it compelling for the next generation of low-latency applications.
Problem

Research questions and friction points this paper is trying to address.

Reducing latency in accelerator access by eliminating host involvement
Enabling direct remote calls to FPGA-based accelerators efficiently
Ensuring performance isolation and scalability for multi-tenant accelerator use
Innovation

Methods, ideas, or system contributions that make the work stand out.

Direct FPGA calls bypassing host involvement
Serverless stateless function abstraction for offloading
Multi-tenancy performance isolation in accelerator calls
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Ziyi Yang
KAUST
K
Krishnan B. Iyer
KAUST
Y
Yixi Chen
KAUST
R
Ran Shi
Microsoft Reserch
Z
Zsolt Istv'an
Technical University of Darmstadt
Marco Canini
Marco Canini
Professor of Computer Science, KAUST
SystemsNetworkingDistributed SystemsMachine Learning
S
Suhaib A. Fahmy
KAUST