Score
Designs, builds, and assesses architectures and operational procedures that keep services available and recoverable across multiple geographic regions, ensuring automated failover, cross‑region replication, routing/traffic management, and runbooks for disaster recovery. Work typically covers replication and consistency strategies, health monitoring and failover automation, data backup/restore across regions, and testing/validation of cross‑region recovery processes.
In cloud environments, selecting optimal data protection strategies for business continuity and disaster recovery remains challenging due to the lack of quantitative foundations for evaluating reliability and aligning with organizational Recovery Time Objectives (RTOs) and operational requirements. Method: This paper proposes an integrated assessment framework that synergistically combines system dynamics modeling and simulation-based optimization. It quantitatively evaluates key performance indicators—including recovery timeliness, data integrity, and system robustness—across public and hybrid cloud scenarios by simulating mainstream recovery mechanisms. Contribution/Results: The framework innovatively applies system dynamics to model time-varying dependencies during recovery processes and establishes interpretable, traceable mappings between policy parameters, technical metrics, and business objectives. Empirical validation demonstrates its reproducibility and practical utility, providing cloud-native organizations with a quantifiable, verifiable, and actionable decision-support methodology for data protection strategy selection.
Azure Cosmos DB struggles to simultaneously achieve fine-grained recovery, low recovery time objective (RTO) and recovery point objective (RPO), and strong consistency under node- to region-level failures. Method: This paper proposes the first partition-level, decentralized cross-region automatic failover architecture. It leverages distributed consensus protocols and partition-granular failure detection and traffic rerouting to fully decentralize metadata coordination and state-machine fault tolerance. Clients can flexibly configure per-partition consistency levels and RPO/RTO targets. Contribution/Results: Experiments demonstrate millisecond-scale RTO for critical partitions, optional RPO = 0 (zero data loss), and robust self-healing across full operational scenarios at scale—supporting over 20 million vCores and 100+ PB of data. This work breaks the conventional region-level disaster recovery paradigm, establishing a new high-availability framework for hyperscale distributed databases.
To address insufficient resilience of complex systems under heterogeneous hardware environments, this paper proposes a fault-adaptive software deployment and redundancy configuration optimization method. We construct a system-level resilience state-space model and introduce a novel equivalence relation to enable quotient-space-based state-space reduction, significantly compressing the state space. Subsequently, we integrate formal model checking with strategy synthesis to automatically derive both an initial deployment configuration and dynamic reconfiguration policies that satisfy multi-level resilience requirements. Our key contributions are: (i) a new equivalence relation enabling efficient, semantics-preserving state-space reduction; and (ii) end-to-end automated synthesis of fault-response and recovery strategies. Experimental evaluation on an autonomous driving system model demonstrates that our approach substantially improves fault recovery latency and system availability, while supporting real-time resilience assurance.
本文提出了一种服务健康工程方法,通过结合遥测、工作流完成情况等手段来检测分布式系统中的静默故障和异步工作停滞问题。
Microservice systems commonly exhibit resilience deficiencies—including localized fault propagation, cascading timeouts, and inconsistent recovery behaviors—yet existing research remains largely descriptive, lacking systematic evidence synthesis and quantitative evaluation. To address this gap, we conduct the first PRISMA-guided systematic literature review (SLR) of 26 high-quality empirical studies published between 2014 and 2025. Our analysis identifies nine core recovery patterns and introduces three novel, empirically grounded artifacts: (1) a reproducible Recovery Pattern Taxonomy; (2) a standardized Resilience Evaluation Score for quantitative assessment; and (3) a constraint-aware decision matrix that explicitly trades off latency, consistency, and cost. Collectively, these contributions establish a structured, quantifiable, and reproducible empirical foundation for resilience-aware microservice design and engineering.
This study addresses the interoperability and migration challenges enterprises face when deploying workloads across AWS and Alibaba Cloud. Through a systematic comparison of architectural designs, service offerings, and operational policies between the two platforms, the research conducts an exploratory case study on migrating IoT workloads using both native and open-source Infrastructure-as-Code (IaC) tools. It reveals critical technical trade-offs inherent in cross-cloud co-deployment for the first time, distills best practices for secure, resilient, and vendor-lock-in-mitigated multicloud deployments, and proposes a multicloud interoperability framework tailored for global enterprises. The findings offer methodological support for empirically grounded multicloud strategies.
Traditional high-availability clusters are often constrained by single points of failure and inefficient resource allocation, making it difficult to meet the continuous availability demands of enterprise-grade systems. This work proposes an integrated High-Availability Cluster (iHAC), which innovatively combines active-active and active-passive architectures to optimize load distribution and failover mechanisms. By harmonizing these approaches, iHAC enhances fault tolerance while significantly improving resource utilization. Simulation experiments conducted using Riverbed Modeler (OPNET) demonstrate that iHAC reduces the average HTTP page response time by over 40%—from 5 seconds to under 3 seconds—compared to conventional solutions. This improvement translates into markedly lower network latency and higher system throughput, underscoring the efficacy of the proposed architecture in real-world deployment scenarios.
This work addresses the limitations of traditional microservice availability assessment, which relies on costly fault injection experiments that are difficult to repeat as architectures evolve and lacks formal characterization of endpoint-level availability. The paper proposes the first runtime availability model grounded in stochastic connectivity, leveraging a typed service dependency graph, replica mappings, and probabilistic measures over node and edge states, combined with a request success predicate, to formally analyze endpoint availability under explicit failures. The approach distinguishes between computational and communication failures, revealing that replica redundancy alone cannot alleviate dependency bottlenecks. It further enables automatic model reconstruction from trace and deployment data for architectural what-if analysis. Experiments demonstrate that the model achieves bounded error relative to theoretical solutions in synthetic scenarios and effectively identifies availability boundaries induced by edge bottlenecks, correlated failures, missing traces, and time-varying faults.
This work addresses the challenge of accurately assessing the resilience of end-to-end applications in real-world communication networks due to limited transparency from network operators. To overcome this, the authors propose DRACO, a novel framework that enables systematic modeling and quantitative evaluation of application deployments across national-scale networks without relying on proprietary operator data. DRACO integrates network modeling, publicly available datasets, synthetic data generation, and resilience metric computation to construct a scalable end-to-end evaluation pipeline. The framework was successfully applied to nationwide networks in Germany and France, demonstrating its effectiveness in handling heterogeneous data sources and enabling large-scale resilience assessments.
本文研究了分布式系统中重试机制在级联故障中的作用,通过引入重试放大因子量化其影响,并提出自适应重试预算方法来优化重试策略。