Score
Designs, implements, and evaluates systems and components that store, retrieve, and manage persistent data—covering file systems, block and object stores, caches, replication/consensus layers, and metadata services—to satisfy requirements for durability, consistency, performance, scalability, and fault tolerance. Works across software and hardware layers (storage protocols, device drivers, SSD/HDD and media characteristics, and distributed/networked architectures) and develops data-layout, caching, replication, backup/recovery, and operational monitoring policies and mechanisms.
Modern data storage systems suffer from latent cross-layer faults due to tight hardware–software coupling across multiple abstraction layers, often leading to silent data corruption or unrecoverable data loss. To address this, we propose the first cross-layer fault-tolerance analysis framework targeting heterogeneous storage stacks—including SSDs, persistent memory, local file systems, and distributed storage. Our approach combines architectural modeling of the full stack, systematic injection of representative defects, and precise tracking of fault propagation across hardware–firmware–software boundaries to expose error propagation paths and consistency violation mechanisms. Through empirical evaluation across widely deployed systems, we identify critical vulnerabilities impacting data integrity and quantify coverage gaps in existing fault-tolerance techniques. The framework provides a scalable, principled methodology for analyzing cross-layer resilience and establishes concrete, actionable directions for designing next-generation highly reliable storage systems.
To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.
This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.
This paper addresses the underappreciated role of persistence in large-memory systems, asserting that “persistence must be treated as a first principle of system design.” It systematically analyzes structural challenges arising from vertical (capacity/latency) and horizontal (network flattening) scaling of the memory hierarchy. The authors propose a full-stack co-designed persistence paradigm: (1) elevating persistence to the foundational axiom of system architecture, and (2) introducing a dual-mechanism persistence model—predictable speculative persistence for performance and strongly consistent deterministic persistence for correctness. By enabling mobile persistence and jointly optimizing memory, interconnect, and storage layers, the approach achieves high throughput, low latency, and cost efficiency simultaneously. Extensive evaluations across diverse workloads demonstrate significant improvements in both system performance and reliability.
Emerging SSD interface standards exacerbate software incompatibility, suboptimal data placement, system instability, and weak cross-platform support. To address these challenges, this paper proposes Reshim—the first fully userspace shim layer for SSDs. Reshim introduces the novel design paradigm of “isolated data placement logic,” enabling dynamic, interface- and application-agnostic rule deployment via host-device co-designed affinity-awareness and data-lifecycle-driven policies—without modifying the OS or applications. It supports zero-intrusion integration with RocksDB, MongoDB, and CacheLib. Evaluation demonstrates 2–6× higher write throughput, up to 6× lower tail latency, and significant write amplification reduction. Reshim matches ZenFS in overall performance while achieving lower latency, greater placement policy flexibility, and broader generality across storage interfaces and workloads.
Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.
Modern storage hierarchies face a fundamental trade-off between load balancing and space efficiency. To address this, we propose Mirror-Optimized Storage Tiering (MOST), a co-design strategy integrating mirroring with tiered storage. MOST implements dynamic hot-data identification and cross-tier mirroring via Cerberus—a user-space storage management layer built atop CacheLib—thereby eliminating the high-overhead data migrations inherent in conventional tiering. Its core innovation lies in employing lightweight mirroring to enhance I/O parallelism and bandwidth utilization while preserving the space efficiency of tiered storage. Experimental evaluation across diverse I/O-intensive and dynamic workloads demonstrates that Cerberus achieves an average 32% throughput improvement over state-of-the-art approaches; gains are especially pronounced in NVMe+SSD hybrid tiers.
In cloud environments, selecting optimal data protection strategies for business continuity and disaster recovery remains challenging due to the lack of quantitative foundations for evaluating reliability and aligning with organizational Recovery Time Objectives (RTOs) and operational requirements. Method: This paper proposes an integrated assessment framework that synergistically combines system dynamics modeling and simulation-based optimization. It quantitatively evaluates key performance indicators—including recovery timeliness, data integrity, and system robustness—across public and hybrid cloud scenarios by simulating mainstream recovery mechanisms. Contribution/Results: The framework innovatively applies system dynamics to model time-varying dependencies during recovery processes and establishes interpretable, traceable mappings between policy parameters, technical metrics, and business objectives. Empirical validation demonstrates its reproducibility and practical utility, providing cloud-native organizations with a quantifiable, verifiable, and actionable decision-support methodology for data protection strategy selection.
This study addresses the practical disparities and co-evolution between high-performance computing (HPC) and edge computing architectures within the cloud continuum. It presents the first large-scale empirical analysis based on 396 real-world, production-grade AWS architectures. Methodologically, we propose a multidimensional, data-driven framework encompassing service topology identification, storage type classification, architectural complexity quantification, and ML service integration statistics. Results reveal systematic differences—and complementary patterns—between HPC and edge architectures across four dimensions: core service composition (e.g., EC2 versus Greengrass/Lambda), storage design paradigms (parallel file systems versus distributed lightweight caches), complexity distributions, and ML embedding strategies. This work delivers the first industry-scale architectural benchmark for the cloud continuum, providing empirically grounded insights and methodological foundations for cross-domain architecture design, resource optimization, and cloud-native convergence of HPC and edge computing.
This work addresses the gap in current systems education, where learning resources often consist of superficial tutorials or AI-generated summaries that inadequately convey foundational design principles and thus fail to cultivate robust engineering capabilities. To remedy this, we propose a structured learning pathway centered on seminal research papers from distributed systems, operating systems, and big data domains. Integrating insights from leading academic curricula and industry practices, our approach emphasizes technical depth and problem-solving reasoning. By engaging learners in close reading of original literature, critical analysis of architectural trade-offs, and cross-domain synthesis, the framework fosters a deep understanding of underlying mechanisms and cultivates systems thinking—thereby equipping practitioners to effectively tackle complex engineering challenges and progress toward professional-level systems expertise.