chaos engineering

Designs, builds, and runs controlled experiments that inject faults, resource perturbations, latency, and configuration changes into running software systems to validate hypotheses about failure modes and to measure and improve system resilience, observability, and recovery behavior. Analyzes experiment outcomes to identify brittle components, refine monitoring and incident response playbooks, and prioritize engineering changes that harden systems against real-world failures.

chaosengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$187K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address insufficient resilience of complex systems under heterogeneous hardware environments, this paper proposes a fault-adaptive software deployment and redundancy configuration optimization method. We construct a system-level resilience state-space model and introduce a novel equivalence relation to enable quotient-space-based state-space reduction, significantly compressing the state space. Subsequently, we integrate formal model checking with strategy synthesis to automatically derive both an initial deployment configuration and dynamic reconfiguration policies that satisfy multi-level resilience requirements. Our key contributions are: (i) a new equivalence relation enabling efficient, semantics-preserving state-space reduction; and (ii) end-to-end automated synthesis of fault-response and recovery strategies. Experimental evaluation on an autonomous driving system model demonstrates that our approach substantially improves fault recovery latency and system availability, while supporting real-time resilience assurance.

Automated framework for resilient complex systems under failuresGenerates resilient initial configurations and reconfiguration policiesOptimized adaptive distribution and replication of software components

Reasonable Experiments in Model-Based Systems Engineering

Sep 12, 2025
JC
Johan Cederbladh
🏛️ Mälardalen University | Eindhoven University of Technology | Stellenbosch University | IT University of Copenhagen | University of Oslo | Universidade Federal Rural de Pernambuco | University of Antwerp

In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.

Deciding if existing experiments can answer new engineering questionsIntelligently reusing experiment-related data to avoid redundant experimentsManaging experimental configuration metadata and results efficiently

Model Discovery and Graph Simulation: A Lightweight Alternative to Chaos Engineering

Jun 12, 2025
AA
Anatoly A. Krasnovsky
🏛️ Innopolis University | QDeep | National Research Tomsk State University

Microservice systems are prone to cascading failures due to strong inter-service dependencies, and conventional chaos engineering relies on costly fault injection in production environments. This paper proposes a lightweight cascading failure prediction method: it automatically constructs a high-fidelity service dependency graph from distributed tracing data and performs Monte Carlo–based stochastic fault propagation simulations on this graph to enable rapid resilience assessment at the design stage—without requiring real-world fault injection. We provide the first theoretical proof that the automatically derived dependency graph supports high-accuracy resilience prediction. Evaluation on a Social Network benchmark shows prediction errors ≤ 0.0004 against ground-truth measurements; mean absolute error (MAE) is 0.025 under no-replica configurations and exactly zero with replicas, demonstrating highly accurate availability estimation capability.

Automate dependency discovery from trace data for failure simulationPredict microservice resilience using lightweight dependency graphsReduce need for full-scale chaos testing with accurate graph models

Continuous Observability Assurance in Cloud-Native Applications

Mar 11, 2025
MC
Maria C. Borges
🏛️ Technische Universität Berlin

In cloud-native microservices, manual and fragmented observability configuration leads to slow fault localization, high resource overhead, and degraded system performance. This paper introduces the first continuous observability assurance methodology, shifting from experience-driven to experiment-driven design. Built upon the Observability eXperimentation (OXN) framework, our approach integrates A/B testing, metric-based feedback loops, and Infrastructure-as-Code (IaC)-enabled automation to dynamically optimize and quantitatively evaluate observability configurations. Evaluated in realistic microservice deployments, our method reduces mean time to detection by 42% on average, decreases sampling overhead by 31%, and—uniquely—enables quantitative validation of how specific observability configurations directly impact Service-Level Objective (SLO) compliance. By establishing a reproducible, iterative, and empirically grounded design paradigm, this work advances observability engineering from ad hoc practice to rigorous, data-driven discipline.

Addressing challenges in fault detection and diagnosis using observability data.Developing a method to guide and automate observability design processes.Ensuring continuous observability in cloud-native microservice applications.

Quantitative Measurement of Cyber Resilience: Modeling and Experimentation

Mar 28, 2023
MJ
Michael J. Weisman
🏛️ DEVCOM Army Research Laboratory | Pennsylvania State University | ICF International | University of California, Irvine

Current cyber-physical systems (CPS) in vehicular environments lack quantitative, experimentally grounded methods for assessing network resilience. Method: This study constructs an experimental testbed replicating real-world truck operational conditions and conducts multiple rounds of malware injection attacks, simultaneously collecting network- and physical-layer data on resistance and recovery behaviors. Contribution/Results: We introduce the novel concept of “bonware” to holistically characterize both cybersecurity defense capability and physical resilience, formalized via an analytically tractable mathematical model. We further define and extract experimentally identifiable, quantitative resilience metrics—termed elastic features—for the first time. Sensitivity analysis confirms these metrics exhibit significant discriminability with respect to attack intensity, defensive strategies, and physical redundancy. This work bridges a critical gap by advancing vehicular CPS resilience from qualitative description to quantifiable, comparable, and optimizable measurement.

Attack RecoveryCyber ResilienceMeasurement Tools

Latest Papers

What's happening recently
View more

This study addresses the lack of consistent verification for faults and low-precision behaviors in LLM-modernized scientific software by proposing a novel differential fault injection framework. The approach systematically evaluates transformation correctness and robustness by injecting identical deterministic faults into both original and LLM-generated code, combined with GAMESS-driven instrumentation and contraction model prediction. Validated through over 2,200 runs, the fault absorption model demonstrates high accuracy and reliability. Furthermore, 200 paired injection experiments yielded fully consistent results, successfully exposing and rectifying deep-seated synchronization defects such as parallel deadlocks and false convergence. These findings effectively ensure the trustworthiness of modernizing legacy scientific codebases using large language models, providing a rigorous methodology for validating behavioral equivalence under faulty conditions.

behavioral equivalencefault injectionlegacy Fortran

This study addresses the lack of systematic understanding regarding the impact of repair loop iteration counts in large language model (LLM)-based software engineering tasks, where prior work often relies on arbitrarily defined repair budgets. Through a cross-task (code generation, test generation, code translation) and cross-model empirical analysis, this work reveals—for the first time—a pronounced diminishing marginal returns phenomenon in iterative repair: performance gains are concentrated within the first 3–4 iterations, with negligible improvements thereafter. The findings underscore that the design of the repair workflow and feedback mechanisms exerts a far greater influence on repair efficacy than the choice of LLM itself. The authors advocate for treating repair budget as a critical experimental variable to ensure reliable, computationally efficient, and reproducible evaluation outcomes.

diminishing returnsiteration limitsLLM-based software engineering

This study addresses the reliability challenges faced by modern web applications due to their inherent complexity and dynamic operating environments. The authors propose a modular self-healing framework grounded in the MAPE-K architecture, which innovatively integrates AutoFix-inspired heuristics with a learning-driven, feedback-guided recovery strategy to enable adaptive fault repair. Evaluated through fault injection experiments and iterative optimization in real-world scenarios, the system achieves an F1 score of 90.7% for fault detection and a 93.2% success rate in recovery, with an average recovery time of just 3.92 seconds. Notably, it sustains throughput at 88%–95% of baseline levels while increasing response time by only 3.1%, thereby significantly enhancing the resilience and autonomous recovery capabilities of web applications.

adaptive recoveryfault toleranceruntime failures

Hot Scholars

NN

Nithin Nagaraj

Complex Systems Programme, National Institute of Advanced Studies, IISc
Complex systemsBrain-inspired machine learningcausality & scientific measures of consciousness
JB

Joan Bruna

Professor of Computer Science, Data Science & Mathematics (aff), Courant Institute and CDS, NYU
Machine Learning
MD

Michael D. Graham

Steenbock Professor of Engineering, Dept. of Chemical and Biological Engineering, UW-Madison
Fluid dynamicscomplex fluids
PR

Philippe Rigollet

Massachusetts Institute of Technology
StatisticsMachine LearningOptimal Transport
AH

Akhila Henry

Amrita Vishwa Vidyapeetham Amritapuri Campus
Machine LearningChaos Theory