NDS: Programmer-Free Offload of High-Performance Near Data Strands

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the strong dependency of near-data processing on specific software-hardware stacks by proposing an automated hardware offloading mechanism for unmodified binaries. By constructing a modular hardware architecture, the system automatically identifies regular loops and decomposes them into parallel execution chains, performing instruction scheduling and data orchestration to achieve transparent near-data computation offloading without programmer intervention, thereby fully decoupling the software stack. Experimental results demonstrate that the proposed system achieves a 3.14× speedup over a single-core host and a 1.82× speedup compared to eight-core parallel execution. Furthermore, performance gains increase significantly with higher invocation frequencies. These findings validate both the feasibility and efficiency of transparently accelerating legacy applications.
📝 Abstract
Near Data Processing (NDP) has the potential to significantly improve system performance and energy by alleviating data movement bottlenecks. However, most NDP proposals pose heavy requirements for the software stack, data layout, and/or the underlying hardware. To broaden NDP adoption, this work focuses on a modular hardware-centric approach that automates the above steps and can target unmodified binaries. This paper presents Near Data Strands (NDS), a framework that orchestrates tasks and data automatically without involving the programmer and the software stack. As a first step towards this ambitious goal, this work focuses on regular loops that meet specific criteria. The framework identifies potential offloadable loops, decomposes them into parallel execution strands, creates a high performance instruction schedule, performs the required data marshaling and orchestration, and initiates near-data execution. We demonstrate that NDS achieves a 3.14x geomean speedup over a single-core host baseline and a 1.82x geomean speedup over 8-core host-parallel variants, all without any modifications to the software stack. These gains grow as repeated loop invocations amortize offload costs, showing that legacy applications can leverage a transparent hardware approach to extract benefits from NDP.
Problem

Research questions and friction points this paper is trying to address.

Near Data Processing
Task Offload
Programmer Transparency
Data Movement Bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

Near Data Processing
Programmer-Free Offload
Near Data Strands
Transparent Hardware
Loop Parallelization
🔎 Similar Papers
No similar papers found.