caching

Designs, implements, and evaluates systems that store copies of data or computation results to reduce latency and load, including cache placement, data structures, eviction and invalidation policies, TTLs, prefetching, and cache hierarchies (in-memory, on-disk, distributed). Builds instrumentation and tests to measure hit/miss rates, staleness, performance and resource trade‑offs, and to ensure correctness, consistency and fault tolerance under concurrency and failures.

caching

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.81
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$208K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

On Configuring a Hierarchy of Storage Media in the Age of NVM

Apr 16, 2018
SG
Shahram Ghandeharizadeh
🏛️ USC | University of California, Irvine

This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.

Determining storage media selection and capacity allocation under budget constraints.Evaluating data replication versus partitioning strategies for performance and recovery.Optimizing memory hierarchy design for caching middleware with NVM and DRAM.

Performance Models for a Two-tiered Storage System

Mar 12, 2025
AS
Aparna Sasidharan
🏛️ IIT | Sandia National Lab | Oak Ridge National Lab

To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.

Design and analyze a two-tiered storage systemDevelop online learning for data tier managementEvaluate performance using queuing and behavioral models

Revisiting Cache Freshness for Emerging Real-Time Applications

Nov 18, 2024
ZM
Ziming Mao
🏛️ UC Berkeley | ICSI

Existing cache freshness mechanisms—particularly traditional TTL—fail to guarantee sub-second data freshness in latency-critical real-time applications, leading to stale cached content and degraded service quality. Method: This paper proposes a lightweight, adaptive, freshness-aware cache refresh strategy. It first systematically identifies the fundamental limitations of TTL for sub-second freshness guarantees, then integrates adaptive control theory, fine-grained cache state monitoring, and latency-sensitive freshness modeling to design a feedback-driven dynamic refresh mechanism. The approach enables online, low-overhead policy adaptation without modifying backend services or incurring additional storage overhead. Contribution/Results: Evaluated under realistic workloads, the strategy reduces P99 freshness error by 62% and cache miss rate by 41%, significantly alleviating the inherent trade-off between timeliness and system resource overhead.

Data FreshnessReal-time ApplicationsTTLs Inefficiency

PUL: Pre-load in Software for Caches Wouldn't Always Play Along

Jun 20, 2025
AB
Arthur Bernhardt
🏛️ Reutlingen University | Technische Universität Darmstadt

Memory latency and bandwidth bottlenecks continue to impede system performance in the post-Moore era. To address this, we present the first systematic demonstration that software prefetching—under near-data processing (NDP) architectures—exhibits superior scalability and efficiency over conventional hardware prefetching. We propose a lightweight preloading paradigm tailored for intelligent memory, which jointly optimizes latency hiding and computational resource utilization via software-driven interleaving of computation and I/O scheduling, coupled with memory-bandwidth-aware load preloading. Experimental evaluation shows that our approach significantly improves compute-unit utilization, and its prefetching efficiency consistently increases with CPU process node advancements. Across multiple generations of hardware platforms, it achieves superior latency-hiding effectiveness compared to state-of-the-art hardware prefetchers.

Address memory latency and bandwidth limitations in system performanceExplore software-based prefetching efficiency in post-Moore systemsOptimize compute utilization via compute/IO interleaving in near-data processing

This work addresses the challenges of efficiency, scalability, and consistency in distributed management of KV caches for large language model (LLM) services. It proposes the first four-dimensional taxonomy—spanning locality, lifetime, ownership, and storage substrate—to systematically analyze over 30 existing studies, identifying five architectural paradigms: local paging, decoupled pipelining, shared storage, memory pooling, and hybrid hierarchical designs. The study reveals that “ownership” is a key differentiator in distributed KV cache architectures and highlights the absence of seven KV-specific metrics in current evaluation methodologies. Furthermore, it connects these gaps to six critical open problems, including fault tolerance, isolation, hierarchical eviction, and speculative decoding.

cache managementdistributed memoryKV cache

Latest Papers

What's happening recently
View more

This work addresses the problem of excessive and ineffective hardware prefetching in datacenter workloads, which wastes precious memory bandwidth. The authors propose a novel hardware-software cooperative prefetching mechanism that leverages page table entries to convey page-level prefetch hints, enabling dynamic control over hardware prefetcher behavior without requiring modifications to the instruction set architecture or application binaries. The approach is compatible with existing state-of-the-art prefetchers—such as BOP, SPP+PPF, and Pythia—and selectively disables prefetching on non-critical pages at runtime to balance coverage and bandwidth efficiency. Experimental evaluation demonstrates that the proposed technique reduces ineffective prefetch requests by approximately 40% under representative datacenter workloads and achieves performance improvements of up to 4.1%.

cache missesdatacenter workloadshardware prefetching

Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.

distributed systemsparallel systemsruntime control

This study addresses the unclear efficacy of prefix cache replacement policies under LLM agent workloads, where complex algorithms frequently underperform. Through production trace analysis, it reveals the structural reasons why LRU proves optimal due to session regularity. The work proposes a compute-saving ratio metric and designs a hybrid lightweight policy combining rapid demotion with compute-aware partial eviction, further optimizing HBM-constrained scenarios via capacity-dependent granularity control. This research validates the effectiveness of recency-based strategies and elucidates the fundamental reasons why sophisticated approaches offer no benefit. Additionally, it open-sources the associated trace datasets and simulator, providing foundational support for future research in this domain.

agentic workloadscache replacementeviction policy

This work addresses the limitations of existing performance evaluation approaches for distributed computing continua, which often focus on a single dimension and fail to holistically characterize the behavior of cross-layer heterogeneous systems. The paper presents the first systematic framework that establishes a comprehensive taxonomy of performance metrics spanning three layers—computation, networking, and application/user—as well as emerging non-functional attributes such as sustainability and observability. By integrating mathematical modeling with cross-layer analysis, the study rigorously defines the applicability, measurement phases, and specifications for each metric category. The resulting framework is both clearly structured and extensible, offering a solid theoretical foundation and practical guidance for unified performance assessment in dynamic, heterogeneous environments.

Cross-layer MetricsDistributed Computing ContinuumHeterogeneous Systems

This study addresses the limitation of the IO500 benchmark, which emphasizes rankings while overlooking system scale, temporal evolution, and log details, by conducting a repository-level feature analysis of 294 submissions. Methodologically transcending conventional aggregated scoring, this work proposes a scale-aware longitudinal analysis framework that integrates descriptive statistics, temporal modeling, and phase-level correlation mining. The analysis reveals strong correlations among analogous I/O phases but weak coupling between bandwidth and metadata performance, alongside notable scale sensitivity. Furthermore, latent performance behavioral characteristics are extracted from execution logs. Ultimately, this research provides critical methodological foundations for leveraging community-submitted data to perform longitudinal performance evaluations of high-performance computing storage systems.

benchmark repositoryhigh-performance computingIO500 benchmark