Score
Designs, implements, and evaluates systems and policies that store and supply frequently used data closer to consumers and that anticipate future accesses, including prefetching algorithms, cache placement, eviction and admission policies, and cache warming. Works across in-memory, multi‑tier, and distributed caches — covering cache coherence and synchronization protocols, low‑latency and build/result caching, cache optimization and management, and tradeoffs among latency, throughput, storage cost, and correctness.
To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.
This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.
This work addresses the challenge of balancing performance and adaptability in distributed caching algorithms under dynamic workloads. We systematically evaluate four mainstream strategies—LRU, LFU, ARC, and TLRU—across microservice and edge cluster architectures, benchmarking them along four dimensions: hit rate, latency, memory overhead, and scalability. To bridge the gap in standardized evaluation, we propose the first unified framework incorporating machine learning–enhanced hybrid caching policies and introduce a load-aware methodology for algorithm selection. Experimental results demonstrate that ML-enhanced strategies improve hit rates by 12–19% and reduce latency by 23% over the best-performing baseline. Furthermore, our methodology enables adaptive strategy selection across heterogeneous deployment scenarios. The framework establishes a reusable, reproducible evaluation paradigm and provides practical design guidelines for distributed caching systems.
Existing KV cache eviction strategies for LLM inference services—particularly generic policies like LRU—suffer from poor adaptability to workload characteristics, leading to suboptimal performance. Method: This paper presents the first empirical analysis of KV caching behavior using real-world cloud service traces, revealing strong skewness, intra-class predictability, and low capacity sensitivity in KV reuse across single- and multi-turn requests. Based on these insights, we propose a workload-aware dynamic eviction policy that jointly models token热度 (access frequency) and时效 (temporal recency), validated via analytical cache modeling, offline trace replay, and online A/B testing. Contribution/Results: On production traces, our approach achieves up to a 23% higher cache hit rate than LRU; under memory-constrained conditions, it reduces end-to-end latency by 18% and increases throughput by over 15%, significantly enhancing deployment efficiency in practical LLM serving systems.
To address the cache adaptability deficiency, poor data locality, and escalating storage/network pressure caused by explosive growth in high-energy physics (HEP) data—e.g., from HL-LHC and DUNE—this paper proposes an hourly adaptive caching strategy. Methodologically, it introduces the first HEP-domain file-granularity hourly access predictor and establishes a two-tier machine learning framework integrating LSTM and CatBoostRegressor to enable fine-grained, dynamic data prefetching and intelligent placement. Evaluated on real SoCal MINI cache traces from August 2024, the approach significantly improves cache hit rate and data locality. Furthermore, the WRENCH simulation platform has been extended to support comprehensive evaluation across multi-level heterogeneous systems. Key contributions include: (1) the first hourly file-access prediction model tailored for HEP workloads; (2) a hybrid ML framework balancing temporal dynamics and feature-rich static attributes; and (3) scalable integration into production-grade simulation infrastructure for realistic system-level assessment.
Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.
This work addresses the challenge of minimizing cache capacity while meeting a specified service-level objective (SLO) hit rate in edge–cloud协同 environments. To this end, the authors propose an SLO-driven dynamic hybrid segmented caching strategy that leverages historical access patterns to adaptively adjust segment proportions within the cache. By dynamically reallocating cache space based on observed workload characteristics, the approach ensures the target hit rate is consistently achieved while substantially reducing both required storage capacity and computational overhead. Experimental evaluation using both real-world and synthetic workload traces demonstrates that, compared to conventional fixed-capacity caching configurations, the proposed strategy effectively lowers cache requirements across diverse load conditions without compromising performance, thereby achieving joint optimization of storage efficiency and system responsiveness.
This study addresses the limitations of static cache replacement and prefetching policies in conventional processors, which struggle to maintain optimal performance across diverse execution phases. For the first time, it systematically evaluates the potential of dynamic policy selection by analyzing 490 execution phases from 49 benchmark programs using the ChampSim simulator. The results demonstrate that static policies incur an average IPC loss of 1.54%, whereas dynamically switching between two carefully selected policies reduces this loss to just 0.11%. Moreover, such dynamic adaptation achieves near-ideal performance in 52.65% of the phases, closely approaching the theoretical upper bound. This work validates the efficacy of dynamic policy switching and establishes a new paradigm for enhancing single-threaded performance.
This work addresses the phase-sensitive interplay between prefetching and replacement policies in modern processors, where static configurations often fail to sustain optimal performance due to dynamic workload behavior. For the first time, it reveals the strong phase dependency of combined L1D/L1I prefetcher and L2 replacement strategies and proposes a lightweight dynamic selection mechanism governed by a single-bit control signal, framing policy switching as an information acquisition problem. Leveraging phase-level performance analysis, execution feedback, passive memory monitoring, and counterfactual evaluation, the approach recovers 62.4%–73.4% of the oracle performance gap across diverse workloads. Notably, the Berti/Gaze combination—dynamically switching only the L1D prefetcher—nearly matches the performance of an eight-policy oracle, achieving an average IPC gap as low as 0.039%.