design cache-aware data structures

Design and implement data structures, memory layouts, caching and memoization mechanisms, and scheduling strategies that optimize data locality and the use of memory hierarchies to minimize cache misses, random I/O, and memory access latency. Measure and analyze memory footprints and access patterns and apply techniques such as tiling, access coalescing, padding, in-memory computation, footprint reduction, and threading/allocation policies to guide placement, allocation, and runtime decisions.

designcache-awaredatastructures

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.49
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address memory-access performance bottlenecks in tree structures on heterogeneous hardware systems—caused by mismatches between tree layouts and hierarchical memory characteristics—this paper proposes a hardware-aware, generic tree node layout method. Our approach introduces: (1) the first unified node reordering strategy explicitly optimized for hardware attributes including latency, bandwidth, and spatial/temporal locality; and (2) a dual-mode triggering mechanism supporting both offline pre-optimization and online dynamic re-optimization, guided by runtime performance monitoring to enable cross-memory-tier adaptive layout adjustments. Experimental evaluation across diverse heterogeneous platforms demonstrates average performance improvements of 95% for offline-optimized layouts and 75% for online-adaptive layouts over conventional approaches. The method exhibits strong generalizability across tree types and hardware configurations, and delivers practical utility for memory-intensive tree-based applications.

Hardware Resource UtilizationMemory Type OptimizationTree Data Structure

This work addresses the growing performance gap between processors and memory caused by irregular, data-dependent memory access patterns in modern applications, which render traditional prefetchers ineffective. The paper proposes the first three-dimensional structured taxonomy that integrates locality type, implementation level, and machine learning (ML) paradigm to systematically survey and multidimensionally compare ML-based prefetching techniques. Guided by the PRISMA framework for literature selection, the study encompasses supervised, unsupervised, and reinforcement learning approaches, analyzing software, hardware, and hybrid architectures under both online and offline training regimes. It reveals key relationships—such as those between model class and the accuracy-overhead Pareto frontier, model complexity and cache hierarchy placement, and runtime adaptability versus model capacity—and delineates performance boundaries of ML prefetchers in terms of storage overhead, latency, generalization capability, and hardware feasibility, thereby offering theoretical guidance for intelligent prefetcher design.

complex workloadsdata-dependent patternsirregular memory access

This work addresses the joint scheduling of general-purpose computation DAGs in multi-processor systems with a two-level memory hierarchy, requiring co-optimization of load balancing, inter-processor communication overhead, and data movement under cache capacity constraints. We identify a fundamental theoretical limitation of conventional decoupled scheduling and memory management strategies: they can incur worst-case linear deviation from optimal performance, underscoring the necessity of tight compute–memory co-optimization. To this end, we propose a unified integer linear programming (ILP) framework that jointly optimizes task scheduling, data placement, and data migration decisions. Experimental evaluation on standard DAG benchmarks demonstrates that our approach consistently outperforms classical decoupled baselines, achieving significant and simultaneous improvements in both execution time and memory efficiency.

Finding optimal solutions using Integer Linear ProgrammingOptimizing parallelization and memory management jointlyScheduling computational DAGs with memory constraints

Heterogeneous Memory Pool Tuning

May 20, 2025
FV
Filip Vaverka
🏛️ IT4Innovations | VSB - Technical University of Ostrava

To address the challenge of optimizing data placement in HBM+DDR heterogeneous memory systems, this paper proposes a lightweight, non-intrusive, application-level memory analysis and control framework. Leveraging a detailed memory subsystem model on Intel Sapphire Rapids platforms—integrated with empirical bandwidth/latency measurements, runtime allocation tracing, and policy injection—the work quantifies, for the first time, the performance inflection point of HBM-DRAM co-utilization: retaining only 60–75% of critical data in HBM achieves 90% of the platform’s peak performance. The approach requires no source-code modification or recompilation. Evaluated across multiple benchmarks, it significantly improves performance for memory-intensive applications while reducing HBM resource consumption by over 30%. This establishes a deployable, fine-grained data placement optimization paradigm for heterogeneous memory systems.

Analyzing and tuning data placement in heterogeneous memory systemsDetermining optimal data allocation ratios for maximizing platform performanceEvaluating performance of HBM and DDR memory subsystems together

On Configuring a Hierarchy of Storage Media in the Age of NVM

Apr 16, 2018
SG
Shahram Ghandeharizadeh
🏛️ USC | University of California, Irvine

This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.

Determining storage media selection and capacity allocation under budget constraints.Evaluating data replication versus partitioning strategies for performance and recovery.Optimizing memory hierarchy design for caching middleware with NVM and DRAM.

Latest Papers

What's happening recently
View more

This work addresses the growing challenge posed by the expanding data footprint of modern applications, which renders memory systems a critical bottleneck for both performance and energy efficiency—constraints that traditional microarchitectures struggle to overcome. To this end, the paper introduces a data-driven microarchitectural design paradigm that systematically integrates lightweight machine learning with application-specific data semantic features across multiple processor components. Key contributions include a reinforcement learning–based hardware prefetcher, a perceptron-driven off-chip access predictor, a synergistic mechanism coordinating prefetching and prediction, and a predictability-aware memory access elimination technique leveraging both address and value repetition. Experimental results demonstrate that the proposed approach substantially outperforms state-of-the-art solutions, delivering significant improvements in both performance and energy efficiency.

data-aware designenergy efficiencymemory bottleneck

This work addresses the performance bottleneck in skip lists caused by cache misses and proposes Foresight, a lightweight and easily integrable cache-friendly optimization. Foresight improves node layout by predicting access patterns and incorporates a tailored concurrency control mechanism to effectively reduce cache misses. The approach is compatible with a wide range of both sequential and concurrent skip list implementations. Experimental results demonstrate that Foresight achieves up to a 45% throughput improvement in microbenchmarks and delivers a 15% end-to-end performance gain in the DBx1000 in-memory database system.

cache missesconcurrent data structuresmemory efficiency

This work addresses the challenges of efficiency, scalability, and consistency in distributed management of KV caches for large language model (LLM) services. It proposes the first four-dimensional taxonomy—spanning locality, lifetime, ownership, and storage substrate—to systematically analyze over 30 existing studies, identifying five architectural paradigms: local paging, decoupled pipelining, shared storage, memory pooling, and hybrid hierarchical designs. The study reveals that “ownership” is a key differentiator in distributed KV cache architectures and highlights the absence of seven KV-specific metrics in current evaluation methodologies. Furthermore, it connects these gaps to six critical open problems, including fault tolerance, isolation, hierarchical eviction, and speculative decoding.

cache managementdistributed memoryKV cache

This study reevaluates the performance benefits of region-based memory allocators in the context of modern hardware and general-purpose memory allocators. Building upon Berger et al.’s seminal work from nearly 25 years ago, we present the first benchmarking effort incorporating large-scale real-world applications such as Clang and Blender. We introduce a novel methodology to quantitatively assess the impact of memory fragmentation on spatial and temporal locality. Combining state-of-the-art allocator implementations, performance profiling tools, and detailed memory access pattern analysis, our experiments demonstrate that region-based allocation continues to significantly enhance memory locality and reduce fragmentation on contemporary systems, thereby improving execution efficiency. These findings not only corroborate but also extend the original conclusions of prior research.

custom memory allocationgeneral-purpose allocatorslocality

This work addresses the performance degradation in large language model (LLM) inference on multi-chiplet NUMA GPUs, where non-uniform memory access and inter-chiplet communication induce kernel latency and poor memory locality. The study presents the first systematic classification of operand sharing patterns across workgroups in LLM kernels—categorized as global, partial, or private—and leverages memory trace analysis, workgroup-level access modeling, and cycle-accurate simulation to demonstrate that each pattern necessitates distinct data placement and scheduling strategies. Building on these insights, the authors propose a sub-group-aware co-scheduling mechanism coupled with an optimized data layout scheme, which significantly enhances execution efficiency and memory locality for LLM kernels on multi-chiplet GPU architectures.

kernel latencyLLMmemory locality

Hot Scholars

MS

Mohammad Sadrosadati

Senior Researcher and Lecturer, ETH Zürich
Heterogeneous ComputingProcessing-In-MemoryMemory SystemsInterconnection Networks
LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
TH

Torsten Hoefler

Professor of Computer Science at ETH Zurich
High Performance ComputingDeep LearningNetworkingMessage Passing Interface
IS

Ion Stoica

Professor of Computer Science, UC Berkeley
Cloud ComputingNetworkingDistributed SystemsBig Data
GP

Gennady Pekhimenko

University of Toronto
Computer ArchitectureSystemsSystems for MLMachine Learning