host‑device data movement optimization

Designs and implements mechanisms and transformations to minimize latency, bandwidth use, and memory overhead when moving data between host (CPU) memory and device/accelerator memory, including buffer management, batching, pipelining, overlap of communication and computation, and data layout or compression techniques. Builds and analyzes runtime and compiler-driven strategies—DMA scheduling, PCIe/NVLink usage, memory pinning, zero‑copy, and asynchronous transfers—to maximize throughput, reduce stalls, and control memory footprint.

host‑devicedatamovementoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work challenges the optimality of traditional DMA for I/O in low-latency, high-concurrency workloads—such as microservices and serverless computing—when deployed over cache-coherent interconnects like CXL 3.0. The authors propose and evaluate “programmable I/O”: a CPU-centric architecture where data movement and control are executed via explicit load/store instructions, eliminating dedicated DMA engines and complex address translation. Key contributions include: (1) the first implementation of open cache-coherence-protocol-enabled device-side cache state awareness on real CXL hardware; (2) a lightweight device state machine and memory-mapping optimization; and (3) native support for fine-grained RPC, streaming operator offloading, and serverless network interfaces. Experiments demonstrate substantially reduced communication latency, throughput competitive with DMA, and superior end-to-end performance across all three target scenarios compared to both conventional DMA and PCIe-based memory-mapped PIO.

Challenges DMA-based I/O efficiency for modern cache-coherent interconnectsDemonstrates coherence protocol advantages over DMA in real hardwareExplores programmed I/O benefits for fine-grained communication workloads

Heterogeneous Memory Pool Tuning

May 20, 2025
FV
Filip Vaverka
🏛️ IT4Innovations | VSB - Technical University of Ostrava

To address the challenge of optimizing data placement in HBM+DDR heterogeneous memory systems, this paper proposes a lightweight, non-intrusive, application-level memory analysis and control framework. Leveraging a detailed memory subsystem model on Intel Sapphire Rapids platforms—integrated with empirical bandwidth/latency measurements, runtime allocation tracing, and policy injection—the work quantifies, for the first time, the performance inflection point of HBM-DRAM co-utilization: retaining only 60–75% of critical data in HBM achieves 90% of the platform’s peak performance. The approach requires no source-code modification or recompilation. Evaluated across multiple benchmarks, it significantly improves performance for memory-intensive applications while reducing HBM resource consumption by over 30%. This establishes a deployable, fine-grained data placement optimization paradigm for heterogeneous memory systems.

Analyzing and tuning data placement in heterogeneous memory systemsDetermining optimal data allocation ratios for maximizing platform performanceEvaluating performance of HBM and DDR memory subsystems together

To address PCIe bandwidth bottlenecks in LLM serving—specifically limiting prefix cache loading and model switching—the paper introduces Multipath Memory Access (MMA), a novel memory access mechanism enabling coordinated multi-path data transfer between GPU and host memory over heterogeneous interconnects (PCIe and NVLink). MMA requires no code modification and deploys transparently via dynamic library injection. By breaking the single-path bandwidth ceiling, MMA achieves a GPU–host memory peak throughput of 245 GB/s (4.62× improvement), reduces first-token latency by 1.14–2.38×, and cuts model-switching latency by 1.12–2.48× under vLLM’s sleeping mode. This work establishes a deployable, low-level memory access paradigm for high-throughput, low-latency LLM inference services.

Addresses PCIe bandwidth bottleneck in LLM servicesEnables multipath data transfer between GPU and host memoryImproves token generation and model switching latency

This work addresses the lack of unified coordination over the full lifecycle of DMA buffers in existing AI data transfer libraries, which undermines safety and performance under high load. To resolve this, we propose dmaplane, the first system that introduces buffer orchestration as a standalone abstraction within the Linux kernel. It provides a unified /dev/dmaplane user-space API to manage allocation, cross-device sharing, secure deallocation, and synchronization. dmaplane integrates NUMA-aware allocation, a kernel-level RDMA engine, direct GPU BAR mapping, and credit-based flow control, substantially enhancing reliability and throughput. Experiments demonstrate that GPU BAR mapping outperforms cudaMemcpy, RDMA WRITE WITH IMMEDIATE enables efficient cross-machine key-value cache transfers, and the system maintains strong safety guarantees with low overhead even under high load.

AI data pathsbuffer orchestrationDMA

A Modern Primer on Processing in Memory

Dec 05, 2020
OM
O. Mutlu
🏛️ ETH Zürich | University Illinois Urbana–Champaign | NVIDIA | MangoBoost Inc.

To address performance, energy-efficiency, and latency bottlenecks arising from data movement between main memory and processors in data-intensive computing, this paper proposes and systematically advances the practical deployment of Processing-in-Memory (PIM). We introduce a unified PIM architecture framework that innovatively integrates 3D-stacked logic layers, on-die in-memory computing units, memory controller–integrated accelerators, on-chip ECC-enhanced DRAM, and hardware-level RowHammer mitigation. This holistic design significantly reduces data movement overhead, achieving energy-efficiency improvements of several-fold to over an order of magnitude across representative workloads, while ensuring scalability and reliability. The framework establishes a foundation for low-power, high-throughput memory-compute infrastructure applicable to both server and mobile platforms. By bridging architectural innovation with system-level implementation, our work accelerates the transition of PIM from theoretical concept to real-world deployment.

Big data processingData movement reductionEnergy-efficient computing

Latest Papers

What's happening recently
View more

This work addresses the performance bottlenecks caused by communication complexity in large-scale heterogeneous chiplet systems by proposing a topology-agnostic dynamic computation migration framework. Instead of merely relocating data, the framework innovatively migrates entire computational contexts—including both code and associated data—to more favorable locations. It integrates a multi-bandwidth-domain chiplet architecture, a hierarchical routing mechanism, and a lightweight machine learning–assisted traffic prediction and scheduling strategy to enable communication-aware load placement and adaptive routing optimization. Experimental results demonstrate migration success rates of 75.2%–97.9%, average latency reductions of 16.4%–62.5%, and up to a 12.5× improvement in throughput. Under large language model (LLM) workloads, the system achieves average improvements of 4.9× in execution time, 5.9× in throughput, and 1.8× in energy efficiency.

chiplet-based systemscommunication complexitycompute relocation

This work addresses the performance degradation in large language model (LLM) inference on multi-chiplet NUMA GPUs, where non-uniform memory access and inter-chiplet communication induce kernel latency and poor memory locality. The study presents the first systematic classification of operand sharing patterns across workgroups in LLM kernels—categorized as global, partial, or private—and leverages memory trace analysis, workgroup-level access modeling, and cycle-accurate simulation to demonstrate that each pattern necessitates distinct data placement and scheduling strategies. Building on these insights, the authors propose a sub-group-aware co-scheduling mechanism coupled with an optimized data layout scheme, which significantly enhances execution efficiency and memory locality for LLM kernels on multi-chiplet GPU architectures.

kernel latencyLLMmemory locality

Traditional virtual memory–assisted buffer management struggles to efficiently support data migration and access in multi-tier memory architectures. This work proposes vmcacheⁿ, a framework that extends the conventional two-level (DRAM–disk) buffering mechanism to an n-level hierarchy (DRAM–remote memory–disk). Leveraging the operating system’s virtual memory subsystem and page migration facilities, vmcacheⁿ constructs a multi-tier cache pool by integrating remote memory technologies such as NUMA and CXL. To enable fine-grained and low-overhead cross-tier page migration, the framework introduces a new system call, move_pages2. Experimental evaluation under the TPC-C workload demonstrates that vmcacheⁿ achieves up to a 4× improvement in query throughput compared to the original vmcache.

buffer managementn-tierpage migration

Transformer inference on hardware accelerators is often bottlenecked not by computational capacity but by paged data movement and interconnect bandwidth. This work proposes a system-accelerator co-design that replaces large on-chip SRAM with small caches and a paged streaming scheduler, enabling explicit overlap of computation and data transfer through a DMA-compute-DMA-out pipeline and 4KB-tiled matrix multiplication on the loosely coupled systolic array MatrixFlow. Evaluated using an extended Gem5-AcceSys full-system simulation framework, the proposed approach achieves up to 22× speedup over a CPU-only baseline and outperforms existing loosely and tightly coupled accelerators by 5–8×. Notably, it attains 80% of the performance achievable with on-chip HBM while operating under standard PCIe host memory constraints.

bandwidth bottleneckdata movementhardware acceleration

Existing near-memory processing (NMP) approaches for dynamic large language model (LLM) serving suffer from inefficiencies due to coarse-grained key-value (KV) cache management and inflexible attention execution. To address these limitations, this work proposes Helios, a hybrid-bonded 3D-DRAM-based LLM serving accelerator that leverages a hardware-software co-design methodology. Helios introduces a spatial-aware KV cache allocation mechanism and customized inter-PE communication primitives to enable efficient execution of distributed block-wise attention. Experimental results demonstrate that Helios significantly improves both performance and energy efficiency: it achieves an average speedup of 3.25× and 3.36× higher energy efficiency compared to state-of-the-art GPU and NMP baselines, while reducing per-token generation latency by up to 72% at P50 and 76% at P99.

attention executiondynamic workloadsKV cache management