implement backups and recovery

Designs, implements, and operates backup and recovery systems and processes, including automated and on‑demand backups, retention and lifecycle policies, encryption, storage placement (including offsite/secondary stores), and replication mechanisms. Builds and verifies recovery and failover procedures, scripts, and tests to restore data and systems to defined recovery point and time objectives while ensuring data integrity and consistency.

implementbackupsandrecovery

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$192K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Modeling and Simulation of Data Protection Systems for Business Continuity and Disaster Recovery

Dec 01, 2025
SN
Sašo Nikolovski
🏛️ AUE -FON University | University "St. Kliment Ohridski"

In cloud environments, selecting optimal data protection strategies for business continuity and disaster recovery remains challenging due to the lack of quantitative foundations for evaluating reliability and aligning with organizational Recovery Time Objectives (RTOs) and operational requirements. Method: This paper proposes an integrated assessment framework that synergistically combines system dynamics modeling and simulation-based optimization. It quantitatively evaluates key performance indicators—including recovery timeliness, data integrity, and system robustness—across public and hybrid cloud scenarios by simulating mainstream recovery mechanisms. Contribution/Results: The framework innovatively applies system dynamics to model time-varying dependencies during recovery processes and establishes interpretable, traceable mappings between policy parameters, technical metrics, and business objectives. Empirical validation demonstrates its reproducibility and practical utility, providing cloud-native organizations with a quantifiable, verifiable, and actionable decision-support methodology for data protection strategy selection.

Comparative analysis of cloud-based recovery solutions for reliabilityModeling and simulation of data protection systems for business continuityProposes a framework for selecting and maintaining organizational recovery solutions

Modern data storage systems suffer from latent cross-layer faults due to tight hardware–software coupling across multiple abstraction layers, often leading to silent data corruption or unrecoverable data loss. To address this, we propose the first cross-layer fault-tolerance analysis framework targeting heterogeneous storage stacks—including SSDs, persistent memory, local file systems, and distributed storage. Our approach combines architectural modeling of the full stack, systematic injection of representative defects, and precise tracking of fault propagation across hardware–firmware–software boundaries to expose error propagation paths and consistency violation mechanisms. Through empirical evaluation across widely deployed systems, we identify critical vulnerabilities impacting data integrity and quantify coverage gaps in existing fault-tolerance techniques. The framework provides a scalable, principled methodology for analyzing cross-layer resilience and establishes concrete, actionable directions for designing next-generation highly reliable storage systems.

Detecting latent bugs across hardware and software layersEnsuring fault tolerance in complex data storage systemsMaintaining data integrity and system recovery capabilities

It's a Feature, Not a Bug: Secure and Auditable State Rollback for Confidential Cloud Applications

Nov 17, 2025
QB
Quinn Burke
🏛️ University of Wisconsin-Madison | Intel Labs

In cloud computing, untrusted storage interfaces are vulnerable to rollback and replay attacks, compromising the integrity of confidential application decisions. Existing hardware-enforced state continuity mechanisms treat all rollbacks as malicious, thus precluding legitimate rollbacks required for fault recovery. This paper introduces Rebound, the first framework enabling fine-grained differentiation between malicious and legitimate rollbacks in confidential cloud environments. Rebound features a policy-authorized reference monitor—hardware-protected via trusted execution—supporting atomic state updates, controlled rollback, and tamper-evident audit logging. Leveraging formal verification, policy-driven access control, and end-to-end logging, we evaluate Rebound in GitLab CI, demonstrating efficient and secure version management for binaries, configurations, and data. The system incurs low overhead while providing strong security guarantees against unauthorized or inconsistent state transitions.

Enable policy-authorized legitimate rollbacks for recoveryPrevent replay and rollback attacks on cloud applicationsProvide secure state transition mediation with audit logs

Defining Atomicity (and Integrity) for Snapshots of Storage in Forensic Computing

May 21, 2025
JO
Jenny Ottmann
🏛️ Friedrich-Alexander-Universität Erlangen-Nürnberg | University of Lausanne

In digital forensics, the atomicity and integrity of storage snapshots lack rigorous definitions that jointly guarantee both instantaneousness and causal ordering—undermining evidentiary admissibility in legal proceedings. To address this, we propose a novel atomicity definition grounded in causal consistency, overcoming the limitation of conventional time-based atomicity models. We further rectify conceptual flaws in existing integrity definitions and introduce a revised, theoretically sound yet engineering-practical integrity criterion—explicitly supporting copy-on-write (CoW) implementations. Our approach integrates causal modeling, formal snapshot semantics, CoW mechanism analysis, and formalization of forensic quality criteria, yielding a verifiable snapshot semantic framework. This work establishes the first theoretical foundation for forensic tool design that unifies causal ordering with instantaneous state capture, thereby significantly enhancing the forensic validity and judicial admissibility of live data acquisition.

Defining atomicity for forensic storage snapshotsEnsuring causality-consistent memory acquisitionFixing integrity issues in existing definitions

Latest Papers

What's happening recently
View more

This study addresses the critical challenge that restoring IT backups alone is insufficient to resume production after ransomware attacks on manufacturing systems, due to deep interdependencies among IT, operational technology (OT), physical processes, identity management, and supply chains. The work reframes recovery as a problem of interdependent continuity in critical infrastructure and introduces, for the first time, the concept of “Minimum Viable Factory Recovery” (MVF Recovery), shifting the objective from full-system restoration to capability-oriented minimal trusted operations. Drawing on a PRISMA-guided multi-source systematic review integrating academic literature, standards, government guidelines, and real-world incidents, the study identifies nine failure modes in recovery efforts and proposes a capability-centered recovery framework. It further establishes an evidence-driven recovery lifecycle model and outlines directions for benchmarking, offering actionable recovery targets for critical manufacturing infrastructure.

critical infrastructuremanufacturing systemsMinimum Viable Factory

This study addresses the absence of governance frameworks for digital legacies after a user’s death and the coordination difficulties faced by bereaved survivors. Through a multidimensional content analysis of 800 Reddit posts, it investigates sociotechnical collaboration and conflict surrounding post-mortem digital privacy and security. The findings reveal the critical role of devices as access gateways and identify data loss as a primary harm. Building on these insights, the authors construct a post-mortem digital governance framework organized along dimensions such as assets and actors. This framework establishes a foundation for mechanism design to support survivor coordination, thereby bridging significant theoretical and practical gaps in the field of digital legacy governance.

Access Control FrameworkContent AnalysisDigital Remains

This work addresses the challenge of achieving exactly-once semantics in distributed systems recovering from crashes under a shared-nothing architecture. By formally modeling the information-theoretic boundaries inherent in crash recovery, it demonstrates that relying solely on locally persisted state inevitably leads to either duplicate or missed deliveries. The paper presents the first rigorous, machine-checkable definition of exactly-once semantics together with precise conditions for its realizability. Leveraging a formal framework built in Isabelle/HOL that integrates information-theoretic limits, crash timing, receiver-side fencing, and evidence validity analysis, the study proves that conventional dual-write protocols suffer from fundamental uncertainty. It then introduces a provably safe recovery mechanism and quantifies how memory constraints for deduplication and source log truncation impact the guaranteed validity window.

crash recoverydual-write problemdurability

Existing workflow persistence frameworks lack precise, machine-verifiable recovery semantic contracts, often resulting in inconsistent post-crash behaviors or violations of their own guarantees. This work proposes RESUME CONTRACT, which formally specifies six core recovery properties and employs TLA+ modeling alongside Verus verification to establish their independence and correctness. Leveraging a deterministic testing framework, the authors empirically evaluate prominent systems across a state space of 7.4 million configurations, uncovering semantic flaws in widely used frameworks such as LangGraph and CrewAI. Guided by these findings, they develop REMIT, a reference implementation that effectively addresses critical issues including fork semantics, recovery validity, and cross-process duplicate consumption, and has already been successfully deployed.

checkpointconformance contracteffect consistency

Hot Scholars

XZ

Xiaonan Zhang

Assistant Professor of Computer Science, Florida State University
Wireless communication and networkEdge AIInternet of Things
QL

Qiyang Li

University of California, Berkeley
Artificial Intelligence in RoboticsMachine Learning
SL

Sergey Levine

UC Berkeley, Physical Intelligence
Machine LearningRoboticsReinforcement Learning
SP

Soujanya Ponnapalli

Postdoc, University of California, Berkeley
Distributed SystemsStorage SystemsDecentralized Systems
ZS

Zachary Sunberg

University of Colorado Boulder
POMDPsArtificial IntelligenceRoboticsAutonomous Vehicles