Score
Designs, builds, and analyzes algorithms and mechanisms that decide which cached entries to remove or invalidate when storage, freshness, or consistency constraints require it; this includes specifying replacement rules, invalidation protocols, TTL/expiry behavior, and prioritization criteria for eviction. The work covers evaluation of eviction strategies against access patterns, cost of recomputation, staleness and coherence requirements, and operational triggers (capacity pressure, updates, time-based expiry) to meet performance and correctness goals.
This work addresses the challenge in hierarchical caching networks where conventional request-time-based eviction policies fail to assess the impact of removing aligned storage blocks on the connectivity of downstream critical services, often causing service disruptions. The authors model aligned eviction as a weighted vertex separation problem on a graph, precisely computing the downstream demand cut of candidate blocks to reject evictions that compromise protected paths and selecting the feasible eviction with minimal impact. They introduce a novel verifiable service-cut certificate mechanism that, for the first time, unifies capacity reclamation, path continuity, and distributed failure within a certifiable interface, and prove that strategies relying solely on historical information can incur unbounded single-step damage. Experiments across 144 scenarios processing 582.9 trillion packets (404.86 PiB) validate theoretical predictions, reveal a zero-impact extremal phase transition point, and enable full impact vectors and audit samples within supplementary material budgets.
Existing cache freshness mechanisms—particularly traditional TTL—fail to guarantee sub-second data freshness in latency-critical real-time applications, leading to stale cached content and degraded service quality. Method: This paper proposes a lightweight, adaptive, freshness-aware cache refresh strategy. It first systematically identifies the fundamental limitations of TTL for sub-second freshness guarantees, then integrates adaptive control theory, fine-grained cache state monitoring, and latency-sensitive freshness modeling to design a feedback-driven dynamic refresh mechanism. The approach enables online, low-overhead policy adaptation without modifying backend services or incurring additional storage overhead. Contribution/Results: Evaluated under realistic workloads, the strategy reduces P99 freshness error by 62% and cache miss rate by 41%, significantly alleviating the inherent trade-off between timeliness and system resource overhead.
This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.
This work addresses the challenges of efficiency, scalability, and consistency in distributed management of KV caches for large language model (LLM) services. It proposes the first four-dimensional taxonomy—spanning locality, lifetime, ownership, and storage substrate—to systematically analyze over 30 existing studies, identifying five architectural paradigms: local paging, decoupled pipelining, shared storage, memory pooling, and hybrid hierarchical designs. The study reveals that “ownership” is a key differentiator in distributed KV cache architectures and highlights the absence of seven KV-specific metrics in current evaluation methodologies. Furthermore, it connects these gaps to six critical open problems, including fault tolerance, isolation, hierarchical eviction, and speculative decoding.
This study addresses the problem of reliable query recovery in semantically transparent caching systems under premise-independent erasures, ensuring that cached objects remain logical consequences of the premise base. The work proposes a logic-derivation-structure-based caching mechanism, introducing strategies of local path interception and shared semantic modules, and establishes the Query Local Projection Theorem and the Residual Leaf Node Law. It presents, for the first time, an exact reliability criterion for semantic caching and proves the optimality of shared modules under homogeneous cost assumptions. By integrating deterministic canonical witnesses, semantic module partitioning, MDS erasure codes, and Datalog modeling, the theoretical analysis shows that cache selection in derivation DAGs of depth two is NP-complete, that MDS-based caching enables recovery of leaf payloads from the loss of at most one packet, and precisely quantifies the relationship among caching overhead, erasure rate, and module cost.