reliability engineering

Designs, builds, and analyzes the architectures, operational practices, tooling, and test frameworks that ensure systems, services, platforms, and data pipelines meet specified availability, durability, latency, and fault‑tolerance objectives under normal and adverse conditions. This work includes defining reliability requirements and SLOs, performing reliability modeling and analysis, conducting resilience and failure‑injection testing, implementing monitoring, alerting and runbooks, and improving production reliability through incident response, postmortems, capacity planning, and automated mitigation.

reliabilityengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.58
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address the challenges of standardizing Site Reliability Engineering (SRE) practices in heterogeneous environments and balancing system reliability with development agility, this paper proposes a customizable SRE process framework. The framework integrates automated operations, multidimensional observability (metrics, logs, traces), error-budget-driven governance, standardized incident response, and progressive delivery (canary and blue-green deployments). It is designed for cross-technology-stack adaptability, enabling contextual implementation of core SRE principles. Evaluated in production systems, the framework reduced mean time to recovery by 42%, decreased unplanned outages by 67%, lowered operational staffing requirements by 35%, and achieved 99.99% service availability. Its primary contribution is the first methodology for customizing SRE processes specifically for heterogeneous environments, empirically demonstrating synergistic improvements in both system reliability and operational efficiency.

Analyzes SRE processes to boost efficiency, reduce downtimeExplores SRE for scalable, reliable software systemsPresents adaptable SRE techniques for diverse environments

This work addresses the unreliability of developer productivity dashboards, which often stems from ad hoc scripts that introduce undetected silent data gaps, eroding organizational trust. To resolve this, we propose a robust ELT pipeline grounded in DAG-based orchestration and the Medallion architecture, decoupling data extraction from transformation to preserve the immutability of raw data. Our approach introduces a state-driven dependency scheduling mechanism and, for the first time, treats metric pipelines as production-grade distributed systems. We emphasize the critical role of immutable raw history in enabling reliable metric redefinition. This methodology significantly enhances data reliability and freshness while effectively eliminating silent failures, thereby restoring organizational confidence in DevOps metrics.

Data ReliabilityDeveloper ProductivityDORA Metrics

本文提出了一种服务健康工程方法,通过结合遥测、工作流完成情况等手段来检测分布式系统中的静默故障和异步工作停滞问题。

Distributed SystemsEnd-to-End User OutcomesReliability

Validating Alerts in Cloud-Native Observability

Oct 27, 2025
MC
Maria C. Borges
🏛️ Technische Universität Berlin

In cloud-native systems, alert rules frequently suffer from false positives and false negatives due to the absence of design-phase validation, while existing tools lack systematic support for alert testing. To address this, we propose the “Alert-as-Experiment” paradigm—the first adaptation of the observability experimentation framework OXN to early-stage alert rule validation. Our approach enables closed-loop, development-time testing and continuous calibration of alert logic via simulated execution, synthetic observation data injection, and real-world scenario replay. It supports parameter tuning and repeatable verification of alert-triggering behavior, shifting alert engineering from empirical practice toward a testable, verifiable, and systematic discipline. Empirical evaluation demonstrates significant reductions in both false positive and false negative rates in production environments, alongside improved fault response latency and system maintainability.

Balancing early fault detection with minimizing false alarmsProviding systematic tools for alert design and testingValidating cloud-native alerts to prevent production outages

This paper addresses the insufficient integration of early-system safety analysis with Model-Based Systems Engineering (MBSE). It comparatively evaluates three functional safety analysis methods—Failure Mode and Effects Analysis (FMEA), Functional Hazard Assessment (FHA), and Fault Feedback and Impact Propagation (FFIP)—and identifies FFIP as superior for detecting emergent behaviors, second-order effects, and fault propagation. Subsequently, it systematically reviews existing MBSE integration practices, categorizing them into four approaches: model transformation, custom algorithm development, built-in toolkits, and manual modeling. The study reveals that current integration efforts are predominantly focused on FMEA, while FHA and FFIP remain in exploratory stages, hindered by the absence of a unified framework and standardized guidelines. To bridge this gap, the paper proposes a novel, full-lifecycle safety analysis integration paradigm aligned with digital engineering transformation—enabling traceable, executable, and evolvable model-driven safety verification.

Analyzing safety analysis methods for early-stage risk identification in complex systemsComparing FMEA FHA and FFIP techniques for modern interconnected systems safetyInvestigating MBSE integration approaches for synergistic lifecycle safety management

Latest Papers

What's happening recently
View more

Automotive electronic control units (ECUs) are intricate systems with hundreds of individual functions, numerous software components, and multiple interdependent tasks. A prevalent structural pattern in these systems are so-called cause-effect chains. While significant research efforts have been dedicated to the temporal analysis and optimization of these chains, particularly minimizing data age and function response times, other crucial non-functional properties remain relatively underexplored. In particular, the safety integrity level (SIL) classification substantially influences the system design by determining task colocation strategies. Improper sharing of functions or interweaving tasks with different safety levels can compromise the integrity of critical functions. Additionally, AUTOSAR basic software (BSW) (e.g. OS, runtime environment, communication stacks, or diagnostics) introduces complexity that varies based on task characteristics and SIL categories. Furthermore, memory requirements present another critical challenge, given the diversity of memory architectures and SIL-specific dependencies that strongly constrain task allocations. This paper thoroughly characterizes a real-world automotive application, describing an automotive application based on SIL constraints, the impact of basic software, and memory requirements. In this context, the Driverator configuration framework is introduced for scalable system analysis.

Automotive ECUAUTOSAR Basic SoftwareMemory Constraints

Traditional high-availability clusters are often constrained by single points of failure and inefficient resource allocation, making it difficult to meet the continuous availability demands of enterprise-grade systems. This work proposes an integrated High-Availability Cluster (iHAC), which innovatively combines active-active and active-passive architectures to optimize load distribution and failover mechanisms. By harmonizing these approaches, iHAC enhances fault tolerance while significantly improving resource utilization. Simulation experiments conducted using Riverbed Modeler (OPNET) demonstrate that iHAC reduces the average HTTP page response time by over 40%—from 5 seconds to under 3 seconds—compared to conventional solutions. This improvement translates into markedly lower network latency and higher system throughput, underscoring the efficacy of the proposed architecture in real-world deployment scenarios.

high-availability clustersresource allocationsingle point of failure

This study addresses the growing challenges traditional failure analysis methods face in the era of advanced packaging technologies—such as chiplets, hybrid bonding, and 3D stacking—by conducting an anonymous global survey of over 100 semiconductor design, packaging, and failure analysis organizations. The findings reveal that 69% of respondents prioritize heterogeneous integration products (mean importance score: 7.92/10), while 54% identify hybrid bonding as the most analytically challenging technique. A strong consensus emerges around the need for standardized data formats, with 83% of participants advocating for unified protocols, and high-resolution non-destructive imaging garners substantial support (mean score: 8.18/10). The research systematically identifies critical pain points in sample preparation and 3D structural inspection, offering empirical insights to guide industry standardization and technological innovation.

advanced packagingchipletdata standardization

Hot Scholars

AG

Angelo Garofalo

University of Bologna, ETH Zurich
HW efficient Machine LearningHeterogeneous Computing ArchitecturesMixed-Criticality Systems
JR

Jaan Raik

Tallinn University of Technology
testverificationfault tolerancereliability
LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
AR

Anna Rumshisky

UMass Lowell / Amazon AGI Foundations
Natural Language ProcessingArtificial IntelligenceDeep LearningMachine Learning
SK

Shahar Kvatinsky

Technion - Israel Institute of Technology
MemristorMemristive systemsVLSIRRAM