Score
Designs, implements, and evaluates systems that store copies of data or computation results to reduce latency and load, including cache placement, data structures, eviction and invalidation policies, TTLs, prefetching, and cache hierarchies (in-memory, on-disk, distributed). Builds instrumentation and tests to measure hit/miss rates, staleness, performance and resource trade‑offs, and to ensure correctness, consistency and fault tolerance under concurrency and failures.
This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.
To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.
Existing cache freshness mechanisms—particularly traditional TTL—fail to guarantee sub-second data freshness in latency-critical real-time applications, leading to stale cached content and degraded service quality. Method: This paper proposes a lightweight, adaptive, freshness-aware cache refresh strategy. It first systematically identifies the fundamental limitations of TTL for sub-second freshness guarantees, then integrates adaptive control theory, fine-grained cache state monitoring, and latency-sensitive freshness modeling to design a feedback-driven dynamic refresh mechanism. The approach enables online, low-overhead policy adaptation without modifying backend services or incurring additional storage overhead. Contribution/Results: Evaluated under realistic workloads, the strategy reduces P99 freshness error by 62% and cache miss rate by 41%, significantly alleviating the inherent trade-off between timeliness and system resource overhead.
Memory latency and bandwidth bottlenecks continue to impede system performance in the post-Moore era. To address this, we present the first systematic demonstration that software prefetching—under near-data processing (NDP) architectures—exhibits superior scalability and efficiency over conventional hardware prefetching. We propose a lightweight preloading paradigm tailored for intelligent memory, which jointly optimizes latency hiding and computational resource utilization via software-driven interleaving of computation and I/O scheduling, coupled with memory-bandwidth-aware load preloading. Experimental evaluation shows that our approach significantly improves compute-unit utilization, and its prefetching efficiency consistently increases with CPU process node advancements. Across multiple generations of hardware platforms, it achieves superior latency-hiding effectiveness compared to state-of-the-art hardware prefetchers.
This work addresses the challenges of efficiency, scalability, and consistency in distributed management of KV caches for large language model (LLM) services. It proposes the first four-dimensional taxonomy—spanning locality, lifetime, ownership, and storage substrate—to systematically analyze over 30 existing studies, identifying five architectural paradigms: local paging, decoupled pipelining, shared storage, memory pooling, and hybrid hierarchical designs. The study reveals that “ownership” is a key differentiator in distributed KV cache architectures and highlights the absence of seven KV-specific metrics in current evaluation methodologies. Furthermore, it connects these gaps to six critical open problems, including fault tolerance, isolation, hierarchical eviction, and speculative decoding.
This work addresses the problem of excessive and ineffective hardware prefetching in datacenter workloads, which wastes precious memory bandwidth. The authors propose a novel hardware-software cooperative prefetching mechanism that leverages page table entries to convey page-level prefetch hints, enabling dynamic control over hardware prefetcher behavior without requiring modifications to the instruction set architecture or application binaries. The approach is compatible with existing state-of-the-art prefetchers—such as BOP, SPP+PPF, and Pythia—and selectively disables prefetching on non-critical pages at runtime to balance coverage and bandwidth efficiency. Experimental evaluation demonstrates that the proposed technique reduces ineffective prefetch requests by approximately 40% under representative datacenter workloads and achieves performance improvements of up to 4.1%.
Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.
This study addresses the unclear efficacy of prefix cache replacement policies under LLM agent workloads, where complex algorithms frequently underperform. Through production trace analysis, it reveals the structural reasons why LRU proves optimal due to session regularity. The work proposes a compute-saving ratio metric and designs a hybrid lightweight policy combining rapid demotion with compute-aware partial eviction, further optimizing HBM-constrained scenarios via capacity-dependent granularity control. This research validates the effectiveness of recency-based strategies and elucidates the fundamental reasons why sophisticated approaches offer no benefit. Additionally, it open-sources the associated trace datasets and simulator, providing foundational support for future research in this domain.
This work addresses the limitations of existing performance evaluation approaches for distributed computing continua, which often focus on a single dimension and fail to holistically characterize the behavior of cross-layer heterogeneous systems. The paper presents the first systematic framework that establishes a comprehensive taxonomy of performance metrics spanning three layers—computation, networking, and application/user—as well as emerging non-functional attributes such as sustainability and observability. By integrating mathematical modeling with cross-layer analysis, the study rigorously defines the applicability, measurement phases, and specifications for each metric category. The resulting framework is both clearly structured and extensible, offering a solid theoretical foundation and practical guidance for unified performance assessment in dynamic, heterogeneous environments.
This study addresses the limitation of the IO500 benchmark, which emphasizes rankings while overlooking system scale, temporal evolution, and log details, by conducting a repository-level feature analysis of 294 submissions. Methodologically transcending conventional aggregated scoring, this work proposes a scale-aware longitudinal analysis framework that integrates descriptive statistics, temporal modeling, and phase-level correlation mining. The analysis reveals strong correlations among analogous I/O phases but weak coupling between bandwidth and metadata performance, alongside notable scale sensitivity. Furthermore, latent performance behavioral characteristics are extracted from execution logs. Ultimately, this research provides critical methodological foundations for leveraging community-submitted data to perform longitudinal performance evaluations of high-performance computing storage systems.