disaster recovery planning

Designs, builds, and documents architectures, processes, and automated runbooks that restore systems, services, and data after failures or disasters — covering backup integration, multi‑region failover, recovery mechanisms, and orchestration to meet defined recovery objectives. Produces recovery strategies, project-level recovery action plans and test plans to validate and operationalize service recovery, including automation and maintainable procedures for coordinated failover and restoration.

disasterrecoveryplanning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.66
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$188K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Modeling and Simulation of Data Protection Systems for Business Continuity and Disaster Recovery

Dec 01, 2025
SN
Sašo Nikolovski
🏛️ AUE -FON University | University "St. Kliment Ohridski"

In cloud environments, selecting optimal data protection strategies for business continuity and disaster recovery remains challenging due to the lack of quantitative foundations for evaluating reliability and aligning with organizational Recovery Time Objectives (RTOs) and operational requirements. Method: This paper proposes an integrated assessment framework that synergistically combines system dynamics modeling and simulation-based optimization. It quantitatively evaluates key performance indicators—including recovery timeliness, data integrity, and system robustness—across public and hybrid cloud scenarios by simulating mainstream recovery mechanisms. Contribution/Results: The framework innovatively applies system dynamics to model time-varying dependencies during recovery processes and establishes interpretable, traceable mappings between policy parameters, technical metrics, and business objectives. Empirical validation demonstrates its reproducibility and practical utility, providing cloud-native organizations with a quantifiable, verifiable, and actionable decision-support methodology for data protection strategy selection.

Comparative analysis of cloud-based recovery solutions for reliabilityModeling and simulation of data protection systems for business continuityProposes a framework for selecting and maintaining organizational recovery solutions

To address insufficient resilience of complex systems under heterogeneous hardware environments, this paper proposes a fault-adaptive software deployment and redundancy configuration optimization method. We construct a system-level resilience state-space model and introduce a novel equivalence relation to enable quotient-space-based state-space reduction, significantly compressing the state space. Subsequently, we integrate formal model checking with strategy synthesis to automatically derive both an initial deployment configuration and dynamic reconfiguration policies that satisfy multi-level resilience requirements. Our key contributions are: (i) a new equivalence relation enabling efficient, semantics-preserving state-space reduction; and (ii) end-to-end automated synthesis of fault-response and recovery strategies. Experimental evaluation on an autonomous driving system model demonstrates that our approach substantially improves fault recovery latency and system availability, while supporting real-time resilience assurance.

Automated framework for resilient complex systems under failuresGenerates resilient initial configurations and reconfiguration policiesOptimized adaptive distribution and replication of software components

Azure Cosmos DB struggles to simultaneously achieve fine-grained recovery, low recovery time objective (RTO) and recovery point objective (RPO), and strong consistency under node- to region-level failures. Method: This paper proposes the first partition-level, decentralized cross-region automatic failover architecture. It leverages distributed consensus protocols and partition-granular failure detection and traffic rerouting to fully decentralize metadata coordination and state-machine fault tolerance. Clients can flexibly configure per-partition consistency levels and RPO/RTO targets. Contribution/Results: Experiments demonstrate millisecond-scale RTO for critical partitions, optional RPO = 0 (zero data loss), and robust self-healing across full operational scenarios at scale—supporting over 20 million vCores and 100+ PB of data. This work breaks the conventional region-level disaster recovery paradigm, establishing a new high-availability framework for hyperscale distributed databases.

Handling diverse faults from node failures to regional outagesImplementing partition-level automatic geo failover in Azure Cosmos DBMinimizing Recovery Time Objective while maintaining consistency

This study addresses the critical challenge that restoring IT backups alone is insufficient to resume production after ransomware attacks on manufacturing systems, due to deep interdependencies among IT, operational technology (OT), physical processes, identity management, and supply chains. The work reframes recovery as a problem of interdependent continuity in critical infrastructure and introduces, for the first time, the concept of “Minimum Viable Factory Recovery” (MVF Recovery), shifting the objective from full-system restoration to capability-oriented minimal trusted operations. Drawing on a PRISMA-guided multi-source systematic review integrating academic literature, standards, government guidelines, and real-world incidents, the study identifies nine failure modes in recovery efforts and proposes a capability-centered recovery framework. It further establishes an evidence-driven recovery lifecycle model and outlines directions for benchmarking, offering actionable recovery targets for critical manufacturing infrastructure.

critical infrastructuremanufacturing systemsMinimum Viable Factory

Latest Papers

What's happening recently
View more

本文提出了一种服务健康工程方法,通过结合遥测、工作流完成情况等手段来检测分布式系统中的静默故障和异步工作停滞问题。

Distributed SystemsEnd-to-End User OutcomesReliability

This study addresses the limitation of existing AI agent benchmarks that conflate task planning with fault recovery, thereby hindering the assessment of operational safety in enterprise environments. To this end, we propose UndoBench, a benchmark introducing a novel evaluation paradigm that decouples task completion from fault recovery. Methodologically, counterfactual paired trials are employed to disentangle task competence from recovery capability, while line-level effect histories and environment state oracles are incorporated for precise verification across diverse domain workflows and multiple recovery strategies. Experimental results demonstrate that although agents achieve a nominal success rate of 83.54%, their conditional recovery success rate drops to merely 46.72%. This discrepancy reveals critical safety vulnerabilities, indicating that naive retry mechanisms can readily trigger repeated external side effects.

AI agentsbenchmark evaluationfault recovery

为解决ERP系统中数据集成和流程监控的碎片化问题,本文提出一种企业流程控制塔,通过集成状态观测、语义翻译、机器学习诊断等方法提升IT团队的工作效率。

Electronic Data Interchange (EDI)Enterprise Resource Planning (ERP)Intermediate Document (IDoc)

Hot Scholars

ZW

Ziwei Wang

School of Electrical and Electronic Engineering, Nanyang Technological University
embodied AIroboticscomputer vision
YL

Yonggen Ling

Tencent Robotics X
SLAMVIOSenor FusionComputer Vision
AB

Abdeslam Boularias

Rutgers University
RoboticsArtificial IntelligenceMachine LearningComputer Vision
CL

Chenxin Li

The Chinese University of Hong Kong
Multimodal LLMAgentWorld Model
HX

Huazhe Xu

Tsinghua University
Embodied AIReinforcement LearningComputer VisionDeep Learning