fault injection testing

Designs and implements controlled fault-injection and mutation testing harnesses and generators—covering input-, message-, and program/AST-level mutations, schema-preserving and round-trip (including intent-guided) mutants, and chaos-style experiments that introduce crashes, missing or failing responses, stale reads, distracting outputs, and semantic edge-case values. Builds and runs failure-simulation campaigns to analyze failure modes and recovery ladders, measure system or agent adaptation and patch correctness, detect inter-module and message-level defects, and drive safety-aware testing and verification.

faultinjectiontesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.91
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$196K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of systematic criteria for determining safe termination timing in autonomous driving simulation testing and the inability of conventional coverage metrics to capture critical failures arising from inter-module interactions. To overcome these limitations, the paper proposes Safety-Aware Mutation Testing (SAMT), which integrates system safety analysis methods—such as System-Theoretic Process Analysis (STPA)—into mutation testing to generate semantically meaningful mutants that reflect real-world hazards at the module interaction level. By leveraging STPA-derived safety rules to guide mutation generation, injecting faults at the message level, and modeling temporal constraints, SAMT enables automated scenario construction and rigorous assessment of test adequacy. This approach provides a theoretical foundation for informed test termination decisions and targeted system remediation.

Autonomous Driving SystemsComponent InteractionMutation Testing

This study addresses the "test-blindness" limitation inherent in traditional mutation testing and existing large language model (LLM) approaches, which often generate redundant or ineffective mutants that fail to precisely expose deficiencies in test suites. To overcome this, we propose a test-aware mutant generation framework that incorporates problem descriptions, reference solutions, and base tests into prompt engineering. This guides LLMs, such as Gemini and the GPT series, to produce non-trivial mutants that pass existing tests yet contain genuine logical errors, thereby shifting from blind fault injection to targeted exploration of test blind spots. Evaluated on the HumanEval and MBPP benchmarks, the proposed method achieves fault detection rates of 87.7% and 79.1%, respectively. These results significantly outperform conventional tools like mutmut and test-blind baselines, establishing a new paradigm for efficient mutation testing.

large language modelsmutation testingtest suite adequacy

Traditional mutation testing struggles to generate mutants that exhibit subtle semantic differences and closely resemble real-world programming faults, thereby limiting test effectiveness. This work proposes a novel approach that, for the first time, integrates large language model–driven round-trip translation between code and natural language intent into mutation testing. By leveraging translation discrepancies and controlled perturbations of intended behavior, the method generates high-quality mutants with nuanced semantic variations. Empirical evaluation on 40 real faulty methods demonstrates that the proposed technique—referred to as RTM—significantly improves fault detection rates using substantially smaller test suites: with only 4 and 30 test cases, RTM detects on average 4× and 1.7× more faults, respectively, than conventional approaches, confirming its efficiency and practicality.

code-to-intent translationfault detectionLLM mistranslation

This study addresses the lack of publicly available mutation testing benchmarks for IEC 61131-3 Structured Text (ST) programs widely used in industrial automation, which has hindered reproducible testing research. The authors present STMutants, the first mutation testing dataset specifically designed for PLC ST programs, incorporating seven mutation operators tailored to industrial control domains. Through a rigorous four-stage pipeline—syntactic transformation, compilation validation, manual equivalence screening (with inter-rater agreement κ = 0.87), and observability filtering—the dataset retains 108 high-quality, non-equivalent mutants. This benchmark facilitates research in automated test generation, fault localization, and AI-assisted quality assurance. Leveraging STMutants, the study evaluates three large language models, achieving mutation detection accuracies of 86.1%, 94.4%, and 86.1%, respectively, with statistical analysis confirming significant performance differences among them.

benchmark datasetindustrial automationmutation testing

An Exploratory Study on Using Large Language Models for Mutation Testing

Jun 14, 2024
BW
Bo Wang
🏛️ Beijing Jiaotong University | University of Luxembourg | King’s College London

Prior work lacks systematic empirical evaluation of large language models’ (LLMs) capability to generate high-quality mutants for mutation testing. Method: This paper presents the first large-scale empirical study, evaluating six open- and closed-source LLMs—including GPT-4 and CodeLlama—using multi-strategy prompt engineering on the Defects4J 2.0 and ConDefects Java benchmarks. Contribution/Results: LLM-generated mutants exhibit significantly higher behavioral similarity to real faults and achieve a 93% fault detection rate—19 percentage points higher than traditional rule-based approaches—along with markedly improved diversity. However, LLMs underperform conventional methods in compilation success rate and in producing fewer equivalent or non-viable mutants. This work establishes the first comprehensive empirical foundation and quality assessment framework for LLM-driven intelligent mutation testing.

Assessing trade-offs in LLM-based mutation quality and costComparing LLM-generated mutants with rule-based approaches' effectivenessEvaluating LLMs' performance in mutation testing comprehensively

Latest Papers

What's happening recently
View more

This work addresses the vulnerability of large language model (LLM) APIs to cascading failures in agent systems caused by erroneous, truncated, or corrupted responses, highlighting the urgent need for robustness evaluation. The authors propose the first chaos engineering framework tailored for agent systems, which enables non-intrusive, runtime fault injection at the API layer via an HTTP proxy without modifying source code. They introduce the first fault taxonomy for agent systems, encompassing crash, omission, and value-type faults in both content and tool-call fields, along with a runtime interception and validation mechanism to ensure effective fault triggering. Evaluations across 65 fault configurations on diverse agent systems and LLMs reveal significant performance degradation—up to a 50-percentage-point drop in pass@1—demonstrating that system design, rather than model capability, primarily governs robustness. Furthermore, existing diagnostic methods achieve less than 56% accuracy, underscoring the critical need for improvement.

Agent SystemsChaos EngineeringFault Injection

Traditional mutation testing, which operates at the syntactic level, struggles to detect semantic defects arising from misunderstandings of program intent. This work proposes an intent-based mutation testing approach that, for the first time, treats programming intent itself as the unit of mutation. Leveraging large language models (LLMs), the method semantically rewrites natural language descriptions of intent and automatically generates executable program mutants. By shifting the focus from syntactic transformations to semantic reinterpretations of intent, this approach produces mutants that are both richer in semantics and more structurally complex. Empirical evaluation on 29 programs shows that 55% of the intent-based mutants are not covered by traditional mutation operators and exhibit significant differences from them in both syntactic and semantic dimensions, thereby substantially enhancing the detection of specification- and behavior-level faults.

fault detectionmutation testingprogram specification

This study addresses the insufficient test coverage of agent frameworks and the low reliability of LLM-dependent code by presenting the first empirical investigation into testing adequacy for such frameworks, alongside a novel technique termed HarnessTester. This approach generates faithful test suites based on explicit agent-framework contracts and integrates static analysis with dynamic execution to substantially enhance both the depth and breadth of testing for LLM-dependent code. Experimental results demonstrate that HarnessTester effectively improves line and branch coverage as well as mutation scores. Furthermore, it identifies 122 real-world bugs, including 88 previously unknown defects, 69 of which have already been confirmed by developers.

Agent HarnessLLM-based AgentsSoftware Reliability

This study addresses the lack of benchmarks and high costs associated with manual validation in semantic modeling evaluation by proposing a mutation testing framework for domain class diagrams. By defining eleven semantic mutation operators to inject controllable defects, we establish an automated and scalable method for assessing evaluator detection capabilities. Experiments demonstrate that results from this automated mutation testing align closely with human evaluations across multiple LLM-based judge configurations. This work pioneers a mutation testing paradigm for class diagram semantic evaluation, effectively validating its reliability as a substitute for manual assessment and providing a standardized verification mechanism for semantic analysis agents.

Domain Class DiagramsGround TruthModel Evaluation

Hot Scholars

MP

Mike Papadakis

Associate professor, University of Luxembourg
Software EngineeringMutation TestingSoftware TestingSoftware Evolution
MH

Mark Harman

Research Scientist at Meta & Professor of Software Engineering at UCL
SBSESoftware TestingEvolutionary ComputationProgram Analysis
JM

Jie M. Zhang

Lecturer (Assistant Professor), King's College London
LLMsSE4MLmachine learning testingmutation testing
SA

Shaukat Ali

Simula Research Laboratory
Quantum Software EngineeringSoftware EngineeringSBSESoftware Testing
HL

Huawei Li

Institute of Computing Technology, Chinese Academy of Sciences
computer engineering