memory hierarchies

Designs, builds, or analyzes layered storage and memory systems that trade off speed, capacity, persistence, and cost (e.g., registers, caches, main memory, and secondary/remote storage) and the mechanisms that move, place, and maintain data across those layers (caching, prefetching, eviction, placement, coherence, and consistency policies). Evaluates and models their behavior and correctness under workloads and access patterns using analytic models, simulation, measurement, and profiling to optimize latency, bandwidth, capacity utilization, and resource cost.

memoryhierarchies

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$196K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Performance Models for a Two-tiered Storage System

Mar 12, 2025
AS
Aparna Sasidharan
🏛️ IIT | Sandia National Lab | Oak Ridge National Lab

To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.

Design and analyze a two-tiered storage systemDevelop online learning for data tier managementEvaluate performance using queuing and behavioral models

On Configuring a Hierarchy of Storage Media in the Age of NVM

Apr 16, 2018
SG
Shahram Ghandeharizadeh
🏛️ USC | University of California, Irvine

This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.

Determining storage media selection and capacity allocation under budget constraints.Evaluating data replication versus partitioning strategies for performance and recovery.Optimizing memory hierarchy design for caching middleware with NVM and DRAM.

Fundamentals of Caching Layered Data objects

Apr 01, 2025
AB
Agrim Bari
🏛️ The University of Texas at Austin | The Pennsylvania State University

This work addresses cache management for hierarchical data objects—such as scalable maps, layered video, VR content, and neural network models—in cloud/edge systems. We propose the first asymptotically exact analytical model for Hierarchical LRU (HLRU), departing from conventional single-layer cache analysis. Our model rigorously characterizes the non-monotonic impact of hierarchy depth, per-layer popularity, and size distribution on cache hit rate. We theoretically prove that “more layers do not necessarily improve performance” and, for the first time, derive the optimal hierarchy depth as a function of the joint popularity–size distribution. This advances cache policy design beyond heuristic approaches by establishing fundamental theoretical bounds and actionable design principles for hierarchical configuration. The framework provides a rigorous analytical foundation for efficient caching of hierarchical data in distributed systems.

Analyzing performance impact of layer count and popularityEvaluating traditional caching policies with layering effectsOptimizing cache management for layered data objects

From Good to Great: Improving Memory Tiering Performance Through Parameter Tuning

Apr 25, 2025
KK
Konstantinos Kanellis
🏛️ University of Wisconsin-Madison

Existing memory tiering systems rely on static thresholds and heuristic policies, rendering them ill-suited to diverse workloads and heterogeneous hardware—leading to suboptimal data placement and migration efficiency. To address this, we propose the first Bayesian optimization–based automated parameter-tuning framework specifically designed for memory tiering. Our method jointly models application runtime behavior and hardware-aware features to enable workload- and hardware-adaptive dynamic configuration of tiering parameters. It operates online atop mainstream tiering systems—including HeMem and HMSDK—to iteratively learn and optimize critical tiering thresholds. Experimental evaluation demonstrates that our approach achieves a 2.0× speedup over default configurations and outperforms the state-of-the-art tiering system by 1.56×. Moreover, it significantly improves cross-workload and cross-platform adaptability, establishing a new foundation for intelligent, self-tuning memory tiering.

Adapting data placement to diverse workloads and hardwareEnhancing performance using Bayesian Optimization for configurationOptimizing memory tiering performance through parameter tuning

Getting the MOST out of your Storage Hierarchy with Mirror-Optimized Storage Tiering

Dec 02, 2025
KT
Kaiwei Tu
🏛️ University of Wisconsin–Madison | Google

Modern storage hierarchies face a fundamental trade-off between load balancing and space efficiency. To address this, we propose Mirror-Optimized Storage Tiering (MOST), a co-design strategy integrating mirroring with tiered storage. MOST implements dynamic hot-data identification and cross-tier mirroring via Cerberus—a user-space storage management layer built atop CacheLib—thereby eliminating the high-overhead data migrations inherent in conventional tiering. Its core innovation lies in employing lightweight mirroring to enhance I/O parallelism and bandwidth utilization while preserving the space efficiency of tiered storage. Experimental evaluation across diverse I/O-intensive and dynamic workloads demonstrates that Cerberus achieves an average 32% throughput improvement over state-of-the-art approaches; gains are especially pronounced in NVMe+SSD hybrid tiers.

Balances load efficiently by mirroring hot data across tiersImproves bandwidth utilization for I/O-intensive workloadsOptimizes storage hierarchies with tiering and mirroring

Latest Papers

What's happening recently
View more

This work proposes an interpretable machine learning–based cross-layer co-design methodology to address the challenge of jointly optimizing reliability and performance in high-density solid-state storage systems. By constructing the first representation learning framework that unifies NAND flash error management with diverse real-world and synthetic workloads—including JEDEC and YCSB benchmarks—and integrating it with an abstracted Flash Translation Layer model, the study systematically analyzes the interaction mechanisms between memory components and firmware algorithms. Extensive experiments across thousands of multi-generation datacenter SSDs demonstrate that the proposed approach significantly enhances storage architecture design efficiency, enables data-driven continuous evolution, and achieves synergistic optimization of both reliability and performance under realistic and synthetic workloads.

Error ManagementInterpretable Machine LearningMemory-Storage Co-Design

This work addresses the challenges of efficiency, scalability, and consistency in distributed management of KV caches for large language model (LLM) services. It proposes the first four-dimensional taxonomy—spanning locality, lifetime, ownership, and storage substrate—to systematically analyze over 30 existing studies, identifying five architectural paradigms: local paging, decoupled pipelining, shared storage, memory pooling, and hybrid hierarchical designs. The study reveals that “ownership” is a key differentiator in distributed KV cache architectures and highlights the absence of seven KV-specific metrics in current evaluation methodologies. Furthermore, it connects these gaps to six critical open problems, including fault tolerance, isolation, hierarchical eviction, and speculative decoding.

cache managementdistributed memoryKV cache

Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.

distributed systemsparallel systemsruntime control

From Minutes to Seconds: Redefining the Five-Minute Rule for AI-Era Memory Hierarchies

Nov 06, 2025
TZ
Tong Zhang
🏛️ ScaleFlux | NVIDIA | Stanford University

This paper revisits and updates the 1987 “Five-Minute Rule” to reflect AI-era memory hierarchies dominated by GPUs and ultra-high-performance SSDs, addressing key limitations of the original rule—namely its neglect of host cost, physical constraints, and workload dynamics. Method: We propose a dynamic resource planning framework integrating DRAM bandwidth/capacity modeling, physics-aware SSD performance modeling, and workload-aware analysis. Leveraging the MQSim-Next simulator, we conduct sensitivity analysis to quantify cache threshold shifts under AI workloads. Contribution/Results: We demonstrate that the DRAM–NAND cache threshold has collapsed from minutes to seconds in AI scenarios—a first-time quantification. We establish NAND flash as a viable *active data layer*, enabling new hardware–software co-design paradigms. Two case studies validate the framework’s ability to expand the system design space, offering principled guidance for GPU-centric storage architecture.

Collapsing DRAM-to-flash threshold from minutes to secondsIntegrating host costs and SSD performance into caching decisionsRedefining the five-minute rule for AI-era memory hierarchies

A Joint Learning Approach to Hardware Caching and Prefetching

Oct 12, 2025
SY
Samuel Yuan
🏛️ The University of Texas at Austin

This work identifies a bidirectional dependency between cache replacement and prefetching policies—a relationship that degrades co-design performance when addressed via independent training. To address this, we propose the first joint learning framework: a shared feature encoder models their mutual interdependence, while contrastive learning enhances representation learning of memory access patterns. Our approach is the first to theoretically establish and empirically exploit this bidirectional coupling, enabling end-to-end co-optimization under dynamic workloads. Experimental evaluation across mainstream benchmarks demonstrates that our method achieves average improvements of 3.2% in cache hit rate and 5.7% in prefetch accuracy over independently trained baselines. These results substantiate a novel paradigm for intelligent, synergistic cache optimization—advancing beyond isolated policy design toward holistic, data-driven co-adaptation of replacement and prefetching.

Addressing suboptimal performance from isolated policy trainingDeveloping shared feature representations for joint learningOptimizing interdependent hardware caching and prefetching policies