durable storage design

Design, build, and analyze persistent storage systems and state stores that ensure data durability across crashes and failures, including on-disk layout and indexing, write paths (e.g., write-ahead logs, flush/sync semantics), recovery and replication protocols, and mechanisms for checkpointing and compaction. Work encompasses specifying and verifying durability and crash-consistency guarantees (atomicity, ordering, idempotence), and evaluating trade-offs among durability, latency, throughput, and resource usage.

durablestoragedesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Modern data storage systems suffer from latent cross-layer faults due to tight hardware–software coupling across multiple abstraction layers, often leading to silent data corruption or unrecoverable data loss. To address this, we propose the first cross-layer fault-tolerance analysis framework targeting heterogeneous storage stacks—including SSDs, persistent memory, local file systems, and distributed storage. Our approach combines architectural modeling of the full stack, systematic injection of representative defects, and precise tracking of fault propagation across hardware–firmware–software boundaries to expose error propagation paths and consistency violation mechanisms. Through empirical evaluation across widely deployed systems, we identify critical vulnerabilities impacting data integrity and quantify coverage gaps in existing fault-tolerance techniques. The framework provides a scalable, principled methodology for analyzing cross-layer resilience and establishes concrete, actionable directions for designing next-generation highly reliable storage systems.

Detecting latent bugs across hardware and software layersEnsuring fault tolerance in complex data storage systemsMaintaining data integrity and system recovery capabilities

The First Principle of Big Memory Systems

Sep 30, 2023
YH
Yu Hua
🏛️ Huazhong University of Science and Technology

This paper addresses the underappreciated role of persistence in large-memory systems, asserting that “persistence must be treated as a first principle of system design.” It systematically analyzes structural challenges arising from vertical (capacity/latency) and horizontal (network flattening) scaling of the memory hierarchy. The authors propose a full-stack co-designed persistence paradigm: (1) elevating persistence to the foundational axiom of system architecture, and (2) introducing a dual-mechanism persistence model—predictable speculative persistence for performance and strongly consistent deterministic persistence for correctness. By enabling mobile persistence and jointly optimizing memory, interconnect, and storage layers, the approach achieves high throughput, low latency, and cost efficiency simultaneously. Extensive evaluations across diverse workloads demonstrate significant improvements in both system performance and reliability.

Achieving cost efficiency and high performance persistenceAnalyzing vertical and horizontal memory hierarchy extensionsExploring flattening networks in traditional storage hierarchies

Arbitration-Free Consistency is Available (and Vice Versa)

Oct 24, 2025
HA
Hagit Attiya
🏛️ Technion — Israel Institute of Technology | Ecole Polytechnique | Institut Polytechnique de Paris | University of Freiburg

Fundamental trade-offs between availability and consistency in distributed storage are often oversimplified in existing work—particularly through extreme models like CAP—lacking precise characterization of how object semantics interact with diverse consistency models. Method: We propose a unified semantic framework that integrates operational semantics with multi-level consistency constraints—including causal consistency, sequential consistency, snapshot isolation, and SQL transactional semantics—covering key data objects such as key-value stores, counters, sets, CRDTs, and transactional databases. Contribution/Results: We introduce the Arbitration-Free Consistency (AFC) theorem, establishing that an object admits highly available implementations if and only if its visibility or read dependencies do not require total-order arbitration. This theorem unifies and generalizes classical results including CAP, providing the first formal, object-centric criterion for designing highly available distributed systems.

Develops semantic framework combining object semantics with consistency modelsIdentifies arbitration-free consistency as enabling available distributed storage implementationsProves arbitration-freedom determines coordination-free consistency across storage systems

On Configuring a Hierarchy of Storage Media in the Age of NVM

Apr 16, 2018
SG
Shahram Ghandeharizadeh
🏛️ USC | University of California, Irvine

This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.

Determining storage media selection and capacity allocation under budget constraints.Evaluating data replication versus partitioning strategies for performance and recovery.Optimizing memory hierarchy design for caching middleware with NVM and DRAM.

Defining Atomicity (and Integrity) for Snapshots of Storage in Forensic Computing

May 21, 2025
JO
Jenny Ottmann
🏛️ Friedrich-Alexander-Universität Erlangen-Nürnberg | University of Lausanne

In digital forensics, the atomicity and integrity of storage snapshots lack rigorous definitions that jointly guarantee both instantaneousness and causal ordering—undermining evidentiary admissibility in legal proceedings. To address this, we propose a novel atomicity definition grounded in causal consistency, overcoming the limitation of conventional time-based atomicity models. We further rectify conceptual flaws in existing integrity definitions and introduce a revised, theoretically sound yet engineering-practical integrity criterion—explicitly supporting copy-on-write (CoW) implementations. Our approach integrates causal modeling, formal snapshot semantics, CoW mechanism analysis, and formalization of forensic quality criteria, yielding a verifiable snapshot semantic framework. This work establishes the first theoretical foundation for forensic tool design that unifies causal ordering with instantaneous state capture, thereby significantly enhancing the forensic validity and judicial admissibility of live data acquisition.

Defining atomicity for forensic storage snapshotsEnsuring causality-consistent memory acquisitionFixing integrity issues in existing definitions

Latest Papers

What's happening recently
View more

This work demonstrates that the atomicity assumptions underpinning traditional Unix tools do not hold under crash consistency, leading to pervasive data corruption and system failures. It introduces the Forward-in-Time Ordering (FITO) analysis framework, which formally proves—for the first time—that no layer in the storage stack, from CPU to persistent media, achieves true atomic persistence. Apparent atomicity arises instead from implicit, cross-layer timing assumptions that leak across abstraction boundaries. System-call-level persistence primitives fail to define a well-defined commit boundary under crashes, resulting in a recursive chain of non-atomic dependencies—a structural flaw inherent to the entire stack. Through formal methods and cross-layer empirical evaluation spanning ext4, NVMe, Linux reboot semantics, and x86-64 architecture, this study identifies FITO assumption violations as the root cause of large-scale cloud outages, database corruptions, wasted AI training cycles, and financial losses, all amplified by retry-induced error propagation.

atomicitycategory mistakecrash consistency

Existing non-volatile memory (NVM)-based key-value stores commonly lack support for modern database interfaces such as snapshots, consistent iterators, and atomic batch operations. To address this limitation, this work proposes and implements FlintKV, a skip-list-based NVM-optimized storage engine that natively supports a full production-grade transactional API while guaranteeing persistent linearizability. FlintKV achieves this through a novel concurrency protocol that integrates multi-version concurrency control with flat combining, leveraging byte-addressable NVM, custom persistence mechanisms, and highly efficient data structures. Experimental results demonstrate that FlintKV significantly enhances performance, achieving up to 75% higher end-to-end throughput compared to state-of-the-art alternatives. The system can be deployed standalone or integrated as a functional enhancement module into existing NVM-based storage systems.

atomic batchdurable linearizabilitykey-value store

Persistent memory programs often sacrifice performance due to excessive use of flush and fence instructions, making it challenging to balance crash consistency with hardware efficiency. This work proposes a black-box binary rewriting technique that requires neither source code nor manual intervention. By leveraging semantic analysis and performance modeling, the method automatically identifies and optimizes redundant persistence instruction sequences while strictly preserving crash consistency semantics. Experimental evaluation on multiple real-world persistent memory applications demonstrates performance improvements of up to 15%, marking the first fully automated and efficient optimization of synchronization instructions for persistent memory programs.

Crash ConsistencyMemory PersistencePerformance Optimization

Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.

distributed systemsparallel systemsruntime control

Hot Scholars

WJ

Weijia Jia

FIEEE, Chair Professor, Beijing Normal University and UIC
Cyber Intelligent ComputingNetworking
ZT

Zhiqing Tang

Associate Professor, Beijing Normal University
Edge ComputingEdge AI SystemsContainerReinforcement Learning
QM

Qianli Ma

Professor of South China University of Technology
Time Series ModellingMachine LearningNatural Language Processing
MP

Manaswini Piduguralla

Indian Institute of Technology Hyderabad
parallel computingdistributed systemsBlockchainMachine leaning
SP

Sathya Peri

Associate Professor, IIT Hyderabad
Parallel and Distributed Systems