high availability

Designs, builds, and analyses services and infrastructure to remain operational and meet uptime/SLA targets despite component failures and maintenance; this includes implementing redundancy, replication, failover, load balancing, automated recovery, monitoring/alerting, graceful degradation, and fault-injection/chaos testing and measuring availability metrics (e.g., MTBF/MTTR) to validate that availability objectives are met.

highavailability

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$194K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address the challenges of standardizing Site Reliability Engineering (SRE) practices in heterogeneous environments and balancing system reliability with development agility, this paper proposes a customizable SRE process framework. The framework integrates automated operations, multidimensional observability (metrics, logs, traces), error-budget-driven governance, standardized incident response, and progressive delivery (canary and blue-green deployments). It is designed for cross-technology-stack adaptability, enabling contextual implementation of core SRE principles. Evaluated in production systems, the framework reduced mean time to recovery by 42%, decreased unplanned outages by 67%, lowered operational staffing requirements by 35%, and achieved 99.99% service availability. Its primary contribution is the first methodology for customizing SRE processes specifically for heterogeneous environments, empirically demonstrating synergistic improvements in both system reliability and operational efficiency.

Analyzes SRE processes to boost efficiency, reduce downtimeExplores SRE for scalable, reliable software systemsPresents adaptable SRE techniques for diverse environments

Assessing Redundancy Strategies to Improve Availability in Virtualized System Architectures

Nov 25, 2025
AS
Alison Silva
🏛️ Universidade de Pernambuco | Universidade Federal Rural de Pernambuco

To address insufficient availability of Nextcloud file servers in private cloud environments, this paper proposes a dual-redundancy architecture jointly operating at the host and virtual machine layers. A system reliability model is constructed using Stochastic Petri Nets (SPNs) and quantitatively evaluated on an Apache CloudStack-based private cloud platform. Compared to single-layer redundancy strategies, the proposed approach significantly improves steady-state system availability, reducing expected downtime by up to 42.6%. The key contributions are: (i) the first formal modeling of cross-layer dual redundancy as a coupled SPN, enabling integrated characterization of inter-layer fault propagation and recovery dynamics; and (ii) a verifiable, reusable modeling framework and decision-support methodology for designing highly available virtualized private cloud infrastructures.

Analyzing private cloud file server availability using Stochastic Petri NetsComparing host and VM redundancy impacts on system downtime reductionEvaluating redundancy strategies for virtualized system availability enhancement

Cloud service reliability assessment lacks empirical validation from diverse, real-world sources, and existing studies fail to characterize cross-layer fault propagation and fault-tolerance efficacy. Method: We construct the first open-source cloud service availability data warehouse, integrating user-reported incidents with operator logs across web services, cloud platforms, and online gaming. We propose a dual-perspective (user-side/operations-side), cross-layer fault data fusion methodology and develop a reproducible simulation framework grounded in real fault traces to quantify performance trade-offs of checkpointing and retry strategies under heterogeneous failure modes. Contribution/Results: Our analysis reveals a counterintuitive yet critical insight: high-level services—due to robust fault-tolerant design—exhibit lower observed failure rates than underlying infrastructure. We publicly release an annotated dataset (GitHub) and analytical tooling, establishing an empirical foundation and methodological framework for cloud reliability modeling and fault-tolerance evaluation.

Comparing failure patterns across different service levelsEvaluating impact of failures on checkpointing and retry mechanismsStandardizing cloud uptime data collection and analysis

Failure Diagnosis in Microservice Systems: A Comprehensive Survey and Analysis

Jun 27, 2024
SZ
Shenglin Zhang
🏛️ Nankai University | Microsoft | Tsinghua University

Microservice systems suffer from low fault localization accuracy and inefficient diagnosis due to strong inter-component interactions and independent deployment, severely undermining system reliability. To address this, we conduct a systematic literature review of 98 publications spanning 2003–2024. Our method integrates qualitative analysis, systematic review, and empirical investigation to construct a standardized diagnostic knowledge graph. We propose a novel multidimensional classification framework covering problem definitions, architectural paradigms, diagnostic dimensions, and evaluation criteria; unify publicly available datasets, toolchains, and metrics; and deliver a reusable technology selection guide. The study establishes the first comprehensive fault diagnosis landscape for microservices, clarifies current bottlenecks and evolutionary trajectories, and enables rapid industrial validation and deployment. Results demonstrate significant improvements in fault localization efficiency and overall system stability.

Error LocalizationMicroservices SystemsReliability

Latest Papers

What's happening recently
View more

本文提出了一种服务健康工程方法,通过结合遥测、工作流完成情况等手段来检测分布式系统中的静默故障和异步工作停滞问题。

Distributed SystemsEnd-to-End User OutcomesReliability

This study addresses the lack of systematic guidance for enterprise software teams in choosing between monolithic and microservices architectures. The work proposes a decision-making framework that integrates technical and organizational factors, evaluating the trade-offs of each architecture across dimensions such as scalability, reliability, deployment efficiency, and organizational complexity. The assessment is grounded in system scale, business requirements, operational maturity, and long-term maintainability. Through architectural pattern analysis, a structured evaluation model, and multiple case studies, the authors develop a practical selection methodology tailored to real-world engineering contexts. This approach offers enterprises clear architectural evolution pathways and actionable guidelines aligned with their developmental stages, thereby significantly enhancing the rationality and sustainability of system design decisions.

MicroservicesMonolithic ArchitectureOrganizational Complexity

This work addresses the limitations of traditional microservice availability assessment, which relies on costly fault injection experiments that are difficult to repeat as architectures evolve and lacks formal characterization of endpoint-level availability. The paper proposes the first runtime availability model grounded in stochastic connectivity, leveraging a typed service dependency graph, replica mappings, and probabilistic measures over node and edge states, combined with a request success predicate, to formally analyze endpoint availability under explicit failures. The approach distinguishes between computational and communication failures, revealing that replica redundancy alone cannot alleviate dependency bottlenecks. It further enables automatic model reconstruction from trace and deployment data for architectural what-if analysis. Experiments demonstrate that the model achieves bounded error relative to theoretical solutions in synthetic scenarios and effectively identifies availability boundaries induced by edge bottlenecks, correlated failures, missing traces, and time-varying faults.

endpoint-level availabilityfault injectionmicroservice availability

This work addresses the degradation of differentiated quality-of-service (QoS) among service tiers (e.g., Premium vs. Freemium) under capacity-constrained failures in replicated databases, where conventional load balancing tends to homogenize performance. The authors propose Priority-aware Load Balancing (PLB), a novel mechanism that incorporates service tiers into post-failure downgrade strategies. PLB dynamically reassigns roles—Premium, Mixed, or Freemium—to healthy replicas via a repair-to-target approach within a shared replica pool, thereby preserving tier-specific QoS guarantees. Implemented as a PostgreSQL JDBC middleware, PLB supports session routing and dynamic role scheduling, effectively combining isolation with resource sharing. Experimental results demonstrate that under single-node and cascading failures, PLB improves median throughput retention for Premium services by 26–28 percentage points, achieves over twice the baseline throughput during the most severe failure phases, and reduces p95 latency by 18.2% compared to round-robin scheduling.

Differentiated QoSFailure HandlingReplicated Databases

This study addresses the lack of a systematic taxonomy for failure modes in multi-provider large language model (LLM) service gateways, which hinders effective detection and diagnosis in production environments. The work proposes the first dual-axis structured failure classification framework for LLM gateways, categorizing failures by origin layer—spanning network/transport, streaming/protocol, state/session, model behavior, and governance/cost—and by detectability (explicit vs. implicit). Through root cause analysis, stress testing, mining of public bug reports, and protocol-level debugging, the authors construct a catalog of five validated failure cases, three of which include reproducible scripts. Notably, the study uncovers two previously undocumented silent failures: conversation history loss due to concurrency races and stream-index collisions corrupting tool-call payloads. These return HTTP 200 responses and pass standard health checks, evading detection without semantic-level observability, thereby posing significant threats to application reliability.

failure modesgateway infrastructuremulti-provider LLM serving