Score
Designs and implements mechanisms and transformations to minimize latency, bandwidth use, and memory overhead when moving data between host (CPU) memory and device/accelerator memory, including buffer management, batching, pipelining, overlap of communication and computation, and data layout or compression techniques. Builds and analyzes runtime and compiler-driven strategies—DMA scheduling, PCIe/NVLink usage, memory pinning, zero‑copy, and asynchronous transfers—to maximize throughput, reduce stalls, and control memory footprint.
To address the compute-memory co-design bottleneck imposed by the “memory wall,” this paper presents a systematic survey of CXL (Compute Express Link)-driven memory-centric interconnect architectures. We propose the first memory-semantics-based three-dimensional taxonomy—encompassing memory pooling, distributed shared memory, and unified memory—to clarify CXL 3.0’s technical pathways toward low latency, strong coherence, and scalability. By integrating address virtualization, MESIF cache coherence, coherent fabric protocols, and cross-node RDMA, our framework enables unified memory semantics across heterogeneous devices. Our analysis reveals CXL’s mechanistic roles in overcoming memory bandwidth limitations, cluster synchronization overhead, and heterogeneity-induced coordination barriers. The work establishes a theoretical foundation for next-generation interconnects and identifies open research directions toward high-bandwidth, low-latency, strongly coherent systems.
This work challenges the optimality of traditional DMA for I/O in low-latency, high-concurrency workloads—such as microservices and serverless computing—when deployed over cache-coherent interconnects like CXL 3.0. The authors propose and evaluate “programmable I/O”: a CPU-centric architecture where data movement and control are executed via explicit load/store instructions, eliminating dedicated DMA engines and complex address translation. Key contributions include: (1) the first implementation of open cache-coherence-protocol-enabled device-side cache state awareness on real CXL hardware; (2) a lightweight device state machine and memory-mapping optimization; and (3) native support for fine-grained RPC, streaming operator offloading, and serverless network interfaces. Experiments demonstrate substantially reduced communication latency, throughput competitive with DMA, and superior end-to-end performance across all three target scenarios compared to both conventional DMA and PCIe-based memory-mapped PIO.
To address the challenge of optimizing data placement in HBM+DDR heterogeneous memory systems, this paper proposes a lightweight, non-intrusive, application-level memory analysis and control framework. Leveraging a detailed memory subsystem model on Intel Sapphire Rapids platforms—integrated with empirical bandwidth/latency measurements, runtime allocation tracing, and policy injection—the work quantifies, for the first time, the performance inflection point of HBM-DRAM co-utilization: retaining only 60–75% of critical data in HBM achieves 90% of the platform’s peak performance. The approach requires no source-code modification or recompilation. Evaluated across multiple benchmarks, it significantly improves performance for memory-intensive applications while reducing HBM resource consumption by over 30%. This establishes a deployable, fine-grained data placement optimization paradigm for heterogeneous memory systems.
To address PCIe bandwidth bottlenecks in LLM serving—specifically limiting prefix cache loading and model switching—the paper introduces Multipath Memory Access (MMA), a novel memory access mechanism enabling coordinated multi-path data transfer between GPU and host memory over heterogeneous interconnects (PCIe and NVLink). MMA requires no code modification and deploys transparently via dynamic library injection. By breaking the single-path bandwidth ceiling, MMA achieves a GPU–host memory peak throughput of 245 GB/s (4.62× improvement), reduces first-token latency by 1.14–2.38×, and cuts model-switching latency by 1.12–2.48× under vLLM’s sleeping mode. This work establishes a deployable, low-level memory access paradigm for high-throughput, low-latency LLM inference services.
This work addresses the lack of unified coordination over the full lifecycle of DMA buffers in existing AI data transfer libraries, which undermines safety and performance under high load. To resolve this, we propose dmaplane, the first system that introduces buffer orchestration as a standalone abstraction within the Linux kernel. It provides a unified /dev/dmaplane user-space API to manage allocation, cross-device sharing, secure deallocation, and synchronization. dmaplane integrates NUMA-aware allocation, a kernel-level RDMA engine, direct GPU BAR mapping, and credit-based flow control, substantially enhancing reliability and throughput. Experiments demonstrate that GPU BAR mapping outperforms cudaMemcpy, RDMA WRITE WITH IMMEDIATE enables efficient cross-machine key-value cache transfers, and the system maintains strong safety guarantees with low overhead even under high load.
To address performance, energy-efficiency, and latency bottlenecks arising from data movement between main memory and processors in data-intensive computing, this paper proposes and systematically advances the practical deployment of Processing-in-Memory (PIM). We introduce a unified PIM architecture framework that innovatively integrates 3D-stacked logic layers, on-die in-memory computing units, memory controller–integrated accelerators, on-chip ECC-enhanced DRAM, and hardware-level RowHammer mitigation. This holistic design significantly reduces data movement overhead, achieving energy-efficiency improvements of several-fold to over an order of magnitude across representative workloads, while ensuring scalability and reliability. The framework establishes a foundation for low-power, high-throughput memory-compute infrastructure applicable to both server and mobile platforms. By bridging architectural innovation with system-level implementation, our work accelerates the transition of PIM from theoretical concept to real-world deployment.
This work addresses the performance bottlenecks caused by communication complexity in large-scale heterogeneous chiplet systems by proposing a topology-agnostic dynamic computation migration framework. Instead of merely relocating data, the framework innovatively migrates entire computational contexts—including both code and associated data—to more favorable locations. It integrates a multi-bandwidth-domain chiplet architecture, a hierarchical routing mechanism, and a lightweight machine learning–assisted traffic prediction and scheduling strategy to enable communication-aware load placement and adaptive routing optimization. Experimental results demonstrate migration success rates of 75.2%–97.9%, average latency reductions of 16.4%–62.5%, and up to a 12.5× improvement in throughput. Under large language model (LLM) workloads, the system achieves average improvements of 4.9× in execution time, 5.9× in throughput, and 1.8× in energy efficiency.
This work addresses the performance degradation in large language model (LLM) inference on multi-chiplet NUMA GPUs, where non-uniform memory access and inter-chiplet communication induce kernel latency and poor memory locality. The study presents the first systematic classification of operand sharing patterns across workgroups in LLM kernels—categorized as global, partial, or private—and leverages memory trace analysis, workgroup-level access modeling, and cycle-accurate simulation to demonstrate that each pattern necessitates distinct data placement and scheduling strategies. Building on these insights, the authors propose a sub-group-aware co-scheduling mechanism coupled with an optimized data layout scheme, which significantly enhances execution efficiency and memory locality for LLM kernels on multi-chiplet GPU architectures.
Traditional virtual memory–assisted buffer management struggles to efficiently support data migration and access in multi-tier memory architectures. This work proposes vmcacheⁿ, a framework that extends the conventional two-level (DRAM–disk) buffering mechanism to an n-level hierarchy (DRAM–remote memory–disk). Leveraging the operating system’s virtual memory subsystem and page migration facilities, vmcacheⁿ constructs a multi-tier cache pool by integrating remote memory technologies such as NUMA and CXL. To enable fine-grained and low-overhead cross-tier page migration, the framework introduces a new system call, move_pages2. Experimental evaluation under the TPC-C workload demonstrates that vmcacheⁿ achieves up to a 4× improvement in query throughput compared to the original vmcache.
Transformer inference on hardware accelerators is often bottlenecked not by computational capacity but by paged data movement and interconnect bandwidth. This work proposes a system-accelerator co-design that replaces large on-chip SRAM with small caches and a paged streaming scheduler, enabling explicit overlap of computation and data transfer through a DMA-compute-DMA-out pipeline and 4KB-tiled matrix multiplication on the loosely coupled systolic array MatrixFlow. Evaluated using an extended Gem5-AcceSys full-system simulation framework, the proposed approach achieves up to 22× speedup over a CPU-only baseline and outperforms existing loosely and tightly coupled accelerators by 5–8×. Notably, it attains 80% of the performance achievable with on-chip HBM while operating under standard PCIe host memory constraints.
Existing near-memory processing (NMP) approaches for dynamic large language model (LLM) serving suffer from inefficiencies due to coarse-grained key-value (KV) cache management and inflexible attention execution. To address these limitations, this work proposes Helios, a hybrid-bonded 3D-DRAM-based LLM serving accelerator that leverages a hardware-software co-design methodology. Helios introduces a spatial-aware KV cache allocation mechanism and customized inter-PE communication primitives to enable efficient execution of distributed block-wise attention. Experimental results demonstrate that Helios significantly improves both performance and energy efficiency: it achieves an average speedup of 3.25× and 3.36× higher energy efficiency compared to state-of-the-art GPU and NMP baselines, while reducing per-token generation latency by up to 72% at P50 and 76% at P99.