Score
Designs, builds, and analyzes system-level mechanisms and operational procedures that detect component or site failures and shift traffic, state, and control to redundant resources with minimal downtime; this includes automated failover orchestration, replication and redundancy architectures, multi- and cross-region failover strategies, failover policies and testing, and the integration of monitoring and rollback controls to ensure reliable switchover and recovery.
To address the challenges of standardizing Site Reliability Engineering (SRE) practices in heterogeneous environments and balancing system reliability with development agility, this paper proposes a customizable SRE process framework. The framework integrates automated operations, multidimensional observability (metrics, logs, traces), error-budget-driven governance, standardized incident response, and progressive delivery (canary and blue-green deployments). It is designed for cross-technology-stack adaptability, enabling contextual implementation of core SRE principles. Evaluated in production systems, the framework reduced mean time to recovery by 42%, decreased unplanned outages by 67%, lowered operational staffing requirements by 35%, and achieved 99.99% service availability. Its primary contribution is the first methodology for customizing SRE processes specifically for heterogeneous environments, empirically demonstrating synergistic improvements in both system reliability and operational efficiency.
To address insufficient resilience of complex systems under heterogeneous hardware environments, this paper proposes a fault-adaptive software deployment and redundancy configuration optimization method. We construct a system-level resilience state-space model and introduce a novel equivalence relation to enable quotient-space-based state-space reduction, significantly compressing the state space. Subsequently, we integrate formal model checking with strategy synthesis to automatically derive both an initial deployment configuration and dynamic reconfiguration policies that satisfy multi-level resilience requirements. Our key contributions are: (i) a new equivalence relation enabling efficient, semantics-preserving state-space reduction; and (ii) end-to-end automated synthesis of fault-response and recovery strategies. Experimental evaluation on an autonomous driving system model demonstrates that our approach substantially improves fault recovery latency and system availability, while supporting real-time resilience assurance.
To address data flow disruptions caused by link failures and congestion in medium-scale enterprise LANs, this paper proposes a lightweight, scalable fault-tolerant switching architecture. We construct a dynamic topology model using NetworkX and simulate link failures and traffic congestion via Scapy. Based on this, we design and implement adaptive fault detection, fast path rerouting, and traffic scheduling algorithms. The system enables controller-free, protocol-agnostic millisecond-scale automatic failover. Experimental evaluation demonstrates that under single-link failure and sudden congestion scenarios, the packet delivery ratio remains above 99.2%, the average recovery time is below 80 ms, and packet loss is reduced by 92% compared to conventional non-fault-tolerant approaches. These results significantly enhance network robustness and service availability.
本文研究了分布式系统中重试机制在级联故障中的作用,通过引入重试放大因子量化其影响,并提出自适应重试预算方法来优化重试策略。
To address insufficient availability of Nextcloud file servers in private cloud environments, this paper proposes a dual-redundancy architecture jointly operating at the host and virtual machine layers. A system reliability model is constructed using Stochastic Petri Nets (SPNs) and quantitatively evaluated on an Apache CloudStack-based private cloud platform. Compared to single-layer redundancy strategies, the proposed approach significantly improves steady-state system availability, reducing expected downtime by up to 42.6%. The key contributions are: (i) the first formal modeling of cross-layer dual redundancy as a coupled SPN, enabling integrated characterization of inter-layer fault propagation and recovery dynamics; and (ii) a verifiable, reusable modeling framework and decision-support methodology for designing highly available virtualized private cloud infrastructures.
This work addresses the limitation of existing benchmarks, which focus solely on accuracy in multi-agent orchestration tasks while neglecting fine-grained diagnosis of failure origins and recovery capabilities. The authors propose a reproducible fault-injection framework to systematically evaluate failure modes, task decomposition quality, and recovery mechanisms within templated enterprise workflows. They introduce two novel metrics: “cascade radius” and failure-mode-specific recovery rates, and employ controlled probes to analyze recovery behavior across different fault types. Experimental results demonstrate that intent-based reasoning routing achieves 100% recovery under adversarial conditions, significantly outperforming keyword-based routing; tool-related failures are fully recoverable, whereas semantic failures prove largely irrecoverable; and cascade radius increases with workflow depth.
This work addresses the degradation of differentiated quality-of-service (QoS) among service tiers (e.g., Premium vs. Freemium) under capacity-constrained failures in replicated databases, where conventional load balancing tends to homogenize performance. The authors propose Priority-aware Load Balancing (PLB), a novel mechanism that incorporates service tiers into post-failure downgrade strategies. PLB dynamically reassigns roles—Premium, Mixed, or Freemium—to healthy replicas via a repair-to-target approach within a shared replica pool, thereby preserving tier-specific QoS guarantees. Implemented as a PostgreSQL JDBC middleware, PLB supports session routing and dynamic role scheduling, effectively combining isolation with resource sharing. Experimental results demonstrate that under single-node and cascading failures, PLB improves median throughput retention for Premium services by 26–28 percentage points, achieves over twice the baseline throughput during the most severe failure phases, and reduces p95 latency by 18.2% compared to round-robin scheduling.
本文提出了一种服务健康工程方法,通过结合遥测、工作流完成情况等手段来检测分布式系统中的静默故障和异步工作停滞问题。
本文介绍了Meta为解决大规模系统持续部署中速度与可靠性之间的矛盾,通过构建服务健康检查器进行自动回滚等方法保障部署安全。
This study addresses the limitation of existing Site Reliability Engineering (SRE) benchmarks, which predominantly evaluate isolated incidents and fail to capture real-world production complexities such as noisy alerts, overlapping failures, and change-driven, long-horizon operations. To bridge this gap, we introduce the first long-horizon, change-driven SRE evaluation paradigm, establishing a continuous operations benchmark tailored for autonomous agents. Leveraging a dual-zone Kubernetes environment with injected concurrent failures, our framework integrates CI/CD pipelines, sealed record bundles, and deterministic offline scoring mechanisms to provide cumulative alerts and persistent workspaces that faithfully replicate production complexity. Experimental results demonstrate that the best-performing method achieves only 41.3 points, revealing that while current agents can effectively correlate and localize faults, executing remediation during active failure windows remains a critical bottleneck.