🤖 AI Summary
To address high-latency accelerator access caused by host involvement, this paper proposes OffRAC—the first host-free remote accelerator direct-call paradigm—elevating hardware accelerators (e.g., FPGAs) to programmable, native first-class computing resources in the network. Methodologically, OffRAC integrates serverless stateless function abstraction with hardware-level multi-tenancy isolation, enabling a lightweight remote invocation protocol, request aggregation mechanism, and dedicated scheduling framework. Its key innovation lies in eliminating the traditional host-mediated bottleneck, thereby achieving low-overhead, strongly isolated, cross-client direct function invocation. Evaluated on a real FPGA platform, the prototype achieves an end-to-end invocation latency of 10.5 μs and throughput of 85 Gbps. OffRAC significantly improves scalability, energy efficiency, and resource utilization in ultra-low-latency data processing scenarios.
📝 Abstract
Modern applications increasingly demand ultra-low latency for data processing, often facilitated by host-controlled accelerators like GPUs and FPGAs. However, significant delays result from host involvement in accessing accelerators. To address this limitation, we introduce a novel paradigm we call Offloading through Remote Accelerator Calls (OffRAC), which elevates accelerators to first-class compute resources. OffRAC enables direct calls to FPGA-based accelerators without host involvement. Utilizing the stateless function abstraction of serverless computing, with applications decomposed into simpler stateless functions, offloading promotes efficient acceleration and distribution of computational loads across the network. To realize this proposal, we present a prototype design and implementation of an OffRAC platform for FPGAs that assembles diverse requests from multiple clients into complete accelerator calls with multi-tenancy performance isolation. This design minimizes the implementation complexity for accelerator users while ensuring isolation and programmability. Results show that the OffRAC approach reduces the latency of network calls to accelerators down to approximately 10.5 us, as well as sustaining high application throughput up to 85Gbps, demonstrating scalability and efficiency, making it compelling for the next generation of low-latency applications.