Score
Designs, builds, or analyzes layered storage and memory systems that trade off speed, capacity, persistence, and cost (e.g., registers, caches, main memory, and secondary/remote storage) and the mechanisms that move, place, and maintain data across those layers (caching, prefetching, eviction, placement, coherence, and consistency policies). Evaluates and models their behavior and correctness under workloads and access patterns using analytic models, simulation, measurement, and profiling to optimize latency, bandwidth, capacity utilization, and resource cost.
To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.
This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.
This work addresses cache management for hierarchical data objects—such as scalable maps, layered video, VR content, and neural network models—in cloud/edge systems. We propose the first asymptotically exact analytical model for Hierarchical LRU (HLRU), departing from conventional single-layer cache analysis. Our model rigorously characterizes the non-monotonic impact of hierarchy depth, per-layer popularity, and size distribution on cache hit rate. We theoretically prove that “more layers do not necessarily improve performance” and, for the first time, derive the optimal hierarchy depth as a function of the joint popularity–size distribution. This advances cache policy design beyond heuristic approaches by establishing fundamental theoretical bounds and actionable design principles for hierarchical configuration. The framework provides a rigorous analytical foundation for efficient caching of hierarchical data in distributed systems.
Existing memory tiering systems rely on static thresholds and heuristic policies, rendering them ill-suited to diverse workloads and heterogeneous hardware—leading to suboptimal data placement and migration efficiency. To address this, we propose the first Bayesian optimization–based automated parameter-tuning framework specifically designed for memory tiering. Our method jointly models application runtime behavior and hardware-aware features to enable workload- and hardware-adaptive dynamic configuration of tiering parameters. It operates online atop mainstream tiering systems—including HeMem and HMSDK—to iteratively learn and optimize critical tiering thresholds. Experimental evaluation demonstrates that our approach achieves a 2.0× speedup over default configurations and outperforms the state-of-the-art tiering system by 1.56×. Moreover, it significantly improves cross-workload and cross-platform adaptability, establishing a new foundation for intelligent, self-tuning memory tiering.
Modern storage hierarchies face a fundamental trade-off between load balancing and space efficiency. To address this, we propose Mirror-Optimized Storage Tiering (MOST), a co-design strategy integrating mirroring with tiered storage. MOST implements dynamic hot-data identification and cross-tier mirroring via Cerberus—a user-space storage management layer built atop CacheLib—thereby eliminating the high-overhead data migrations inherent in conventional tiering. Its core innovation lies in employing lightweight mirroring to enhance I/O parallelism and bandwidth utilization while preserving the space efficiency of tiered storage. Experimental evaluation across diverse I/O-intensive and dynamic workloads demonstrates that Cerberus achieves an average 32% throughput improvement over state-of-the-art approaches; gains are especially pronounced in NVMe+SSD hybrid tiers.
This work proposes an interpretable machine learning–based cross-layer co-design methodology to address the challenge of jointly optimizing reliability and performance in high-density solid-state storage systems. By constructing the first representation learning framework that unifies NAND flash error management with diverse real-world and synthetic workloads—including JEDEC and YCSB benchmarks—and integrating it with an abstracted Flash Translation Layer model, the study systematically analyzes the interaction mechanisms between memory components and firmware algorithms. Extensive experiments across thousands of multi-generation datacenter SSDs demonstrate that the proposed approach significantly enhances storage architecture design efficiency, enables data-driven continuous evolution, and achieves synergistic optimization of both reliability and performance under realistic and synthetic workloads.
This work addresses the challenges of efficiency, scalability, and consistency in distributed management of KV caches for large language model (LLM) services. It proposes the first four-dimensional taxonomy—spanning locality, lifetime, ownership, and storage substrate—to systematically analyze over 30 existing studies, identifying five architectural paradigms: local paging, decoupled pipelining, shared storage, memory pooling, and hybrid hierarchical designs. The study reveals that “ownership” is a key differentiator in distributed KV cache architectures and highlights the absence of seven KV-specific metrics in current evaluation methodologies. Furthermore, it connects these gaps to six critical open problems, including fault tolerance, isolation, hierarchical eviction, and speculative decoding.
Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.
This paper revisits and updates the 1987 “Five-Minute Rule” to reflect AI-era memory hierarchies dominated by GPUs and ultra-high-performance SSDs, addressing key limitations of the original rule—namely its neglect of host cost, physical constraints, and workload dynamics. Method: We propose a dynamic resource planning framework integrating DRAM bandwidth/capacity modeling, physics-aware SSD performance modeling, and workload-aware analysis. Leveraging the MQSim-Next simulator, we conduct sensitivity analysis to quantify cache threshold shifts under AI workloads. Contribution/Results: We demonstrate that the DRAM–NAND cache threshold has collapsed from minutes to seconds in AI scenarios—a first-time quantification. We establish NAND flash as a viable *active data layer*, enabling new hardware–software co-design paradigms. Two case studies validate the framework’s ability to expand the system design space, offering principled guidance for GPU-centric storage architecture.
This work identifies a bidirectional dependency between cache replacement and prefetching policies—a relationship that degrades co-design performance when addressed via independent training. To address this, we propose the first joint learning framework: a shared feature encoder models their mutual interdependence, while contrastive learning enhances representation learning of memory access patterns. Our approach is the first to theoretically establish and empirically exploit this bidirectional coupling, enabling end-to-end co-optimization under dynamic workloads. Experimental evaluation across mainstream benchmarks demonstrates that our method achieves average improvements of 3.2% in cache hit rate and 5.7% in prefetch accuracy over independently trained baselines. These results substantiate a novel paradigm for intelligent, synergistic cache optimization—advancing beyond isolated policy design toward holistic, data-driven co-adaptation of replacement and prefetching.