The DMA Streaming Framework: Kernel-Level Buffer Orchestration for High-Performance AI Data Paths

📅 2026-02-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of unified coordination over the full lifecycle of DMA buffers in existing AI data transfer libraries, which undermines safety and performance under high load. To resolve this, we propose dmaplane, the first system that introduces buffer orchestration as a standalone abstraction within the Linux kernel. It provides a unified /dev/dmaplane user-space API to manage allocation, cross-device sharing, secure deallocation, and synchronization. dmaplane integrates NUMA-aware allocation, a kernel-level RDMA engine, direct GPU BAR mapping, and credit-based flow control, substantially enhancing reliability and throughput. Experiments demonstrate that GPU BAR mapping outperforms cudaMemcpy, RDMA WRITE WITH IMMEDIATE enables efficient cross-machine key-value cache transfers, and the system maintains strong safety guarantees with low overhead even under high load.

Technology Category

Machine Learning: Hardware-aware MLHumans and AI: Human-Aware Planning and Behavior PredictionData Mining & Knowledge Management: Scalability, Parallel & Distributed Systems

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Data management and stream processing for Web, mobile and wireless applicationsSecurity and Privacy: Data transparency and provenanceResponsible Web: Data and user privacy-enhancing technologies for the Web
📝 Abstract
AI transport libraries move bytes efficiently, but they commonly assume that buffers are already correctly allocated, placed, shared, registered, and safe under completion and teardown pressure. This paper presents dmaplane, a Linux kernel module that makes this missing layer explicit as buffer orchestration. dmaplane exposes a stable kernel UAPI via /dev/dmaplane and composes ring-based command channels, DMA buffer lifecycle management, dma-buf export for cross-device sharing, a kernel-space RDMA engine, NUMA-aware allocation and verification, credit-based flow control, low-overhead observability, and GPU memory integration via PCIe BAR pinning. We evaluate orchestration sensitivity with measurements of NUMA cross-node penalties at DRAM scale, completion-safe flow control under sustained RDMA load, and GPU BAR mapping tiers versus cudaMemcpy. We also demonstrate end-to-end disaggregated inference by transferring KV-cache chunks between two machines using RDMA WRITE WITH IMMEDIATE and reconstructing tensor views on the receiver. RDMA measurements use Soft-RoCE; we distinguish measured results from provider-independent properties by construction.
Problem

Research questions and friction points this paper is trying to address.

buffer orchestration
AI data paths
DMA
RDMA
kernel-level management
Innovation

Methods, ideas, or system contributions that make the work stand out.

buffer orchestration
DMA streaming
kernel-level RDMA
NUMA-aware allocation
GPU memory integration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Marco Graziano
Graziano Labs Corp.