Score
Designs, builds, and documents architectures, processes, and automated runbooks that restore systems, services, and data after failures or disasters — covering backup integration, multi‑region failover, recovery mechanisms, and orchestration to meet defined recovery objectives. Produces recovery strategies, project-level recovery action plans and test plans to validate and operationalize service recovery, including automation and maintainable procedures for coordinated failover and restoration.
In cloud environments, selecting optimal data protection strategies for business continuity and disaster recovery remains challenging due to the lack of quantitative foundations for evaluating reliability and aligning with organizational Recovery Time Objectives (RTOs) and operational requirements. Method: This paper proposes an integrated assessment framework that synergistically combines system dynamics modeling and simulation-based optimization. It quantitatively evaluates key performance indicators—including recovery timeliness, data integrity, and system robustness—across public and hybrid cloud scenarios by simulating mainstream recovery mechanisms. Contribution/Results: The framework innovatively applies system dynamics to model time-varying dependencies during recovery processes and establishes interpretable, traceable mappings between policy parameters, technical metrics, and business objectives. Empirical validation demonstrates its reproducibility and practical utility, providing cloud-native organizations with a quantifiable, verifiable, and actionable decision-support methodology for data protection strategy selection.
本文针对DevOps部署中的自动故障恢复难题,提出了一种结合基于规则和机器学习的混合框架,通过实时监控、识别、分类并恢复故障,有效提高了系统的恢复效率与运行稳定性。
To address insufficient resilience of complex systems under heterogeneous hardware environments, this paper proposes a fault-adaptive software deployment and redundancy configuration optimization method. We construct a system-level resilience state-space model and introduce a novel equivalence relation to enable quotient-space-based state-space reduction, significantly compressing the state space. Subsequently, we integrate formal model checking with strategy synthesis to automatically derive both an initial deployment configuration and dynamic reconfiguration policies that satisfy multi-level resilience requirements. Our key contributions are: (i) a new equivalence relation enabling efficient, semantics-preserving state-space reduction; and (ii) end-to-end automated synthesis of fault-response and recovery strategies. Experimental evaluation on an autonomous driving system model demonstrates that our approach substantially improves fault recovery latency and system availability, while supporting real-time resilience assurance.
Azure Cosmos DB struggles to simultaneously achieve fine-grained recovery, low recovery time objective (RTO) and recovery point objective (RPO), and strong consistency under node- to region-level failures. Method: This paper proposes the first partition-level, decentralized cross-region automatic failover architecture. It leverages distributed consensus protocols and partition-granular failure detection and traffic rerouting to fully decentralize metadata coordination and state-machine fault tolerance. Clients can flexibly configure per-partition consistency levels and RPO/RTO targets. Contribution/Results: Experiments demonstrate millisecond-scale RTO for critical partitions, optional RPO = 0 (zero data loss), and robust self-healing across full operational scenarios at scale—supporting over 20 million vCores and 100+ PB of data. This work breaks the conventional region-level disaster recovery paradigm, establishing a new high-availability framework for hyperscale distributed databases.
This study addresses the critical challenge that restoring IT backups alone is insufficient to resume production after ransomware attacks on manufacturing systems, due to deep interdependencies among IT, operational technology (OT), physical processes, identity management, and supply chains. The work reframes recovery as a problem of interdependent continuity in critical infrastructure and introduces, for the first time, the concept of “Minimum Viable Factory Recovery” (MVF Recovery), shifting the objective from full-system restoration to capability-oriented minimal trusted operations. Drawing on a PRISMA-guided multi-source systematic review integrating academic literature, standards, government guidelines, and real-world incidents, the study identifies nine failure modes in recovery efforts and proposes a capability-centered recovery framework. It further establishes an evidence-driven recovery lifecycle model and outlines directions for benchmarking, offering actionable recovery targets for critical manufacturing infrastructure.
本文提出了一种服务健康工程方法,通过结合遥测、工作流完成情况等手段来检测分布式系统中的静默故障和异步工作停滞问题。
This study addresses the limitation of existing AI agent benchmarks that conflate task planning with fault recovery, thereby hindering the assessment of operational safety in enterprise environments. To this end, we propose UndoBench, a benchmark introducing a novel evaluation paradigm that decouples task completion from fault recovery. Methodologically, counterfactual paired trials are employed to disentangle task competence from recovery capability, while line-level effect histories and environment state oracles are incorporated for precise verification across diverse domain workflows and multiple recovery strategies. Experimental results demonstrate that although agents achieve a nominal success rate of 83.54%, their conditional recovery success rate drops to merely 46.72%. This discrepancy reveals critical safety vulnerabilities, indicating that naive retry mechanisms can readily trigger repeated external side effects.
研究了AI代理系统中检查点和回滚(C/R)的安全性问题,通过建立执行模型识别出五种基本故障模式,并展示了它们对安全的影响。
为解决ERP系统中数据集成和流程监控的碎片化问题,本文提出一种企业流程控制塔,通过集成状态观测、语义翻译、机器学习诊断等方法提升IT团队的工作效率。
本文为小型研究软件团队提供了一份简短的灾难规划和恢复指南,以应对IT灾难和其他突发事件。