Score
Design and implement data structures, memory layouts, caching and memoization mechanisms, and scheduling strategies that optimize data locality and the use of memory hierarchies to minimize cache misses, random I/O, and memory access latency. Measure and analyze memory footprints and access patterns and apply techniques such as tiling, access coalescing, padding, in-memory computation, footprint reduction, and threading/allocation policies to guide placement, allocation, and runtime decisions.
To address memory-access performance bottlenecks in tree structures on heterogeneous hardware systems—caused by mismatches between tree layouts and hierarchical memory characteristics—this paper proposes a hardware-aware, generic tree node layout method. Our approach introduces: (1) the first unified node reordering strategy explicitly optimized for hardware attributes including latency, bandwidth, and spatial/temporal locality; and (2) a dual-mode triggering mechanism supporting both offline pre-optimization and online dynamic re-optimization, guided by runtime performance monitoring to enable cross-memory-tier adaptive layout adjustments. Experimental evaluation across diverse heterogeneous platforms demonstrates average performance improvements of 95% for offline-optimized layouts and 75% for online-adaptive layouts over conventional approaches. The method exhibits strong generalizability across tree types and hardware configurations, and delivers practical utility for memory-intensive tree-based applications.
This work addresses the growing performance gap between processors and memory caused by irregular, data-dependent memory access patterns in modern applications, which render traditional prefetchers ineffective. The paper proposes the first three-dimensional structured taxonomy that integrates locality type, implementation level, and machine learning (ML) paradigm to systematically survey and multidimensionally compare ML-based prefetching techniques. Guided by the PRISMA framework for literature selection, the study encompasses supervised, unsupervised, and reinforcement learning approaches, analyzing software, hardware, and hybrid architectures under both online and offline training regimes. It reveals key relationships—such as those between model class and the accuracy-overhead Pareto frontier, model complexity and cache hierarchy placement, and runtime adaptability versus model capacity—and delineates performance boundaries of ML prefetchers in terms of storage overhead, latency, generalization capability, and hardware feasibility, thereby offering theoretical guidance for intelligent prefetcher design.
This work addresses the joint scheduling of general-purpose computation DAGs in multi-processor systems with a two-level memory hierarchy, requiring co-optimization of load balancing, inter-processor communication overhead, and data movement under cache capacity constraints. We identify a fundamental theoretical limitation of conventional decoupled scheduling and memory management strategies: they can incur worst-case linear deviation from optimal performance, underscoring the necessity of tight compute–memory co-optimization. To this end, we propose a unified integer linear programming (ILP) framework that jointly optimizes task scheduling, data placement, and data migration decisions. Experimental evaluation on standard DAG benchmarks demonstrates that our approach consistently outperforms classical decoupled baselines, achieving significant and simultaneous improvements in both execution time and memory efficiency.
To address the challenge of optimizing data placement in HBM+DDR heterogeneous memory systems, this paper proposes a lightweight, non-intrusive, application-level memory analysis and control framework. Leveraging a detailed memory subsystem model on Intel Sapphire Rapids platforms—integrated with empirical bandwidth/latency measurements, runtime allocation tracing, and policy injection—the work quantifies, for the first time, the performance inflection point of HBM-DRAM co-utilization: retaining only 60–75% of critical data in HBM achieves 90% of the platform’s peak performance. The approach requires no source-code modification or recompilation. Evaluated across multiple benchmarks, it significantly improves performance for memory-intensive applications while reducing HBM resource consumption by over 30%. This establishes a deployable, fine-grained data placement optimization paradigm for heterogeneous memory systems.
This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.
This work addresses the growing challenge posed by the expanding data footprint of modern applications, which renders memory systems a critical bottleneck for both performance and energy efficiency—constraints that traditional microarchitectures struggle to overcome. To this end, the paper introduces a data-driven microarchitectural design paradigm that systematically integrates lightweight machine learning with application-specific data semantic features across multiple processor components. Key contributions include a reinforcement learning–based hardware prefetcher, a perceptron-driven off-chip access predictor, a synergistic mechanism coordinating prefetching and prediction, and a predictability-aware memory access elimination technique leveraging both address and value repetition. Experimental results demonstrate that the proposed approach substantially outperforms state-of-the-art solutions, delivering significant improvements in both performance and energy efficiency.
This work addresses the performance bottleneck in skip lists caused by cache misses and proposes Foresight, a lightweight and easily integrable cache-friendly optimization. Foresight improves node layout by predicting access patterns and incorporates a tailored concurrency control mechanism to effectively reduce cache misses. The approach is compatible with a wide range of both sequential and concurrent skip list implementations. Experimental results demonstrate that Foresight achieves up to a 45% throughput improvement in microbenchmarks and delivers a 15% end-to-end performance gain in the DBx1000 in-memory database system.
This work addresses the challenges of efficiency, scalability, and consistency in distributed management of KV caches for large language model (LLM) services. It proposes the first four-dimensional taxonomy—spanning locality, lifetime, ownership, and storage substrate—to systematically analyze over 30 existing studies, identifying five architectural paradigms: local paging, decoupled pipelining, shared storage, memory pooling, and hybrid hierarchical designs. The study reveals that “ownership” is a key differentiator in distributed KV cache architectures and highlights the absence of seven KV-specific metrics in current evaluation methodologies. Furthermore, it connects these gaps to six critical open problems, including fault tolerance, isolation, hierarchical eviction, and speculative decoding.
This study reevaluates the performance benefits of region-based memory allocators in the context of modern hardware and general-purpose memory allocators. Building upon Berger et al.’s seminal work from nearly 25 years ago, we present the first benchmarking effort incorporating large-scale real-world applications such as Clang and Blender. We introduce a novel methodology to quantitatively assess the impact of memory fragmentation on spatial and temporal locality. Combining state-of-the-art allocator implementations, performance profiling tools, and detailed memory access pattern analysis, our experiments demonstrate that region-based allocation continues to significantly enhance memory locality and reduce fragmentation on contemporary systems, thereby improving execution efficiency. These findings not only corroborate but also extend the original conclusions of prior research.
This work addresses the performance degradation in large language model (LLM) inference on multi-chiplet NUMA GPUs, where non-uniform memory access and inter-chiplet communication induce kernel latency and poor memory locality. The study presents the first systematic classification of operand sharing patterns across workgroups in LLM kernels—categorized as global, partial, or private—and leverages memory trace analysis, workgroup-level access modeling, and cycle-accurate simulation to demonstrate that each pattern necessitates distinct data placement and scheduling strategies. Building on these insights, the authors propose a sub-group-aware co-scheduling mechanism coupled with an optimized data layout scheme, which significantly enhances execution efficiency and memory locality for LLM kernels on multi-chiplet GPU architectures.