data caching and prefetching

Designs, implements, and evaluates systems and policies that store and supply frequently used data closer to consumers and that anticipate future accesses, including prefetching algorithms, cache placement, eviction and admission policies, and cache warming. Works across in-memory, multi‑tier, and distributed caches — covering cache coherence and synchronization protocols, low‑latency and build/result caching, cache optimization and management, and tradeoffs among latency, throughput, storage cost, and correctness.

datacachingandprefetching

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.63
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$208K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Performance Models for a Two-tiered Storage System

Mar 12, 2025
AS
Aparna Sasidharan
🏛️ IIT | Sandia National Lab | Oak Ridge National Lab

To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.

Design and analyze a two-tiered storage systemDevelop online learning for data tier managementEvaluate performance using queuing and behavioral models

On Configuring a Hierarchy of Storage Media in the Age of NVM

Apr 16, 2018
SG
Shahram Ghandeharizadeh
🏛️ USC | University of California, Irvine

This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.

Determining storage media selection and capacity allocation under budget constraints.Evaluating data replication versus partitioning strategies for performance and recovery.Optimizing memory hierarchy design for caching middleware with NVM and DRAM.

This work addresses the challenge of balancing performance and adaptability in distributed caching algorithms under dynamic workloads. We systematically evaluate four mainstream strategies—LRU, LFU, ARC, and TLRU—across microservice and edge cluster architectures, benchmarking them along four dimensions: hit rate, latency, memory overhead, and scalability. To bridge the gap in standardized evaluation, we propose the first unified framework incorporating machine learning–enhanced hybrid caching policies and introduce a load-aware methodology for algorithm selection. Experimental results demonstrate that ML-enhanced strategies improve hit rates by 12–19% and reduce latency by 23% over the best-performing baseline. Furthermore, our methodology enables adaptive strategy selection across heterogeneous deployment scenarios. The framework establishes a reusable, reproducible evaluation paradigm and provides practical design guidelines for distributed caching systems.

Compares distributed caching algorithms' performance metricsEvaluates LRU, LFU, ARC, TLRU on hit ratio, latency, scalabilityRecommends hybrid ML-based algorithms for dynamic environments

KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider

Jun 03, 2025
JW
Jiahao Wang
🏛️ Shanghai Jiao Tong University | Alibaba Group

Existing KV cache eviction strategies for LLM inference services—particularly generic policies like LRU—suffer from poor adaptability to workload characteristics, leading to suboptimal performance. Method: This paper presents the first empirical analysis of KV caching behavior using real-world cloud service traces, revealing strong skewness, intra-class predictability, and low capacity sensitivity in KV reuse across single- and multi-turn requests. Based on these insights, we propose a workload-aware dynamic eviction policy that jointly models token热度 (access frequency) and时效 (temporal recency), validated via analytical cache modeling, offline trace replay, and online A/B testing. Contribution/Results: On production traces, our approach achieves up to a 23% higher cache hit rate than LRU; under memory-constrained conditions, it reduces end-to-end latency by 18% and increases throughput by over 15%, significantly enhancing deployment efficiency in practical LLM serving systems.

Characterizing KV cache workload patterns in large-scale LLM servicesImproving LLM serving performance with limited cache capacityOptimizing cache eviction policies for diverse request categories

ML-based Adaptive Prefetching and Data Placement for US HEP Systems

Mar 08, 2025
VS
Venkat Sai Suman Lamba Karanam
🏛️ University of Nebraska-Lincoln

To address the cache adaptability deficiency, poor data locality, and escalating storage/network pressure caused by explosive growth in high-energy physics (HEP) data—e.g., from HL-LHC and DUNE—this paper proposes an hourly adaptive caching strategy. Methodologically, it introduces the first HEP-domain file-granularity hourly access predictor and establishes a two-tier machine learning framework integrating LSTM and CatBoostRegressor to enable fine-grained, dynamic data prefetching and intelligent placement. Evaluated on real SoCal MINI cache traces from August 2024, the approach significantly improves cache hit rate and data locality. Furthermore, the WRENCH simulation platform has been extended to support comprehensive evaluation across multi-level heterogeneous systems. Key contributions include: (1) the first hourly file-access prediction model tailored for HEP workloads; (2) a hybrid ML framework balancing temporal dynamics and feature-rich static attributes; and (3) scalable integration into production-grade simulation infrastructure for realistic system-level assessment.

Adaptive caching strategies for US HEP systems.Exponential data growth challenges storage and compute.ML-based hourly cache and file-level access prediction.

Latest Papers

What's happening recently
View more

Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.

distributed systemsparallel systemsruntime control

This work addresses the challenge of minimizing cache capacity while meeting a specified service-level objective (SLO) hit rate in edge–cloud协同 environments. To this end, the authors propose an SLO-driven dynamic hybrid segmented caching strategy that leverages historical access patterns to adaptively adjust segment proportions within the cache. By dynamically reallocating cache space based on observed workload characteristics, the approach ensures the target hit rate is consistently achieved while substantially reducing both required storage capacity and computational overhead. Experimental evaluation using both real-world and synthetic workload traces demonstrates that, compared to conventional fixed-capacity caching configurations, the proposed strategy effectively lowers cache requirements across diverse load conditions without compromising performance, thereby achieving joint optimization of storage efficiency and system responsiveness.

cache provisioningcomputational overheadedge-cloud

This study addresses the limitations of static cache replacement and prefetching policies in conventional processors, which struggle to maintain optimal performance across diverse execution phases. For the first time, it systematically evaluates the potential of dynamic policy selection by analyzing 490 execution phases from 49 benchmark programs using the ChampSim simulator. The results demonstrate that static policies incur an average IPC loss of 1.54%, whereas dynamically switching between two carefully selected policies reduces this loss to just 0.11%. Moreover, such dynamic adaptation achieves near-ideal performance in 52.65% of the phases, closely approaching the theoretical upper bound. This work validates the efficacy of dynamic policy switching and establishes a new paradigm for enhancing single-threaded performance.

cache replacementdynamic policy selectionout-of-order pipeline

This work addresses the phase-sensitive interplay between prefetching and replacement policies in modern processors, where static configurations often fail to sustain optimal performance due to dynamic workload behavior. For the first time, it reveals the strong phase dependency of combined L1D/L1I prefetcher and L2 replacement strategies and proposes a lightweight dynamic selection mechanism governed by a single-bit control signal, framing policy switching as an information acquisition problem. Leveraging phase-level performance analysis, execution feedback, passive memory monitoring, and counterfactual evaluation, the approach recovers 62.4%–73.4% of the oracle performance gap across diverse workloads. Notably, the Berti/Gaze combination—dynamically switching only the L1D prefetcher—nearly matches the performance of an eight-policy oracle, achieving an average IPC gap as low as 0.039%.

dynamic selectionmicroarchitectural policiesperformance optimization

Hot Scholars

LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
GC

Giuseppe Caire

Professor, Technical University of Berlin, Germany, and Professor of Electrical Engineering (on
Information TheoryCommunicationsSignal ProcessingStatistics
MC

Minquan Cheng

Guangxi Normal University
Coding TheoryCombinatoricsInformation Theory
CZ

Chang Zou

Intern at EPIC Lab, Shanghai Jiao Tong University
Generative modelsImages and Videos generation
GC

Guihai Chen

Professor of Computer Science
Computer Science and Technology