Beyond Rule-Based Mutation Testing: Test-Aware Mutant Generation Using Large Language Models

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the "test-blindness" limitation inherent in traditional mutation testing and existing large language model (LLM) approaches, which often generate redundant or ineffective mutants that fail to precisely expose deficiencies in test suites. To overcome this, we propose a test-aware mutant generation framework that incorporates problem descriptions, reference solutions, and base tests into prompt engineering. This guides LLMs, such as Gemini and the GPT series, to produce non-trivial mutants that pass existing tests yet contain genuine logical errors, thereby shifting from blind fault injection to targeted exploration of test blind spots. Evaluated on the HumanEval and MBPP benchmarks, the proposed method achieves fault detection rates of 87.7% and 79.1%, respectively. These results significantly outperform conventional tools like mutmut and test-blind baselines, establishing a new paradigm for efficient mutation testing.
📝 Abstract
Mutation testing evaluates test-suite adequacy by injecting synthetic faults into program code. However, traditional rule-based tools often generate large numbers of trivial, redundant, or equivalent mutants that limit their practical use for identifying gaps in a test suite. While recent large language model (LLM)-based approaches generate more realistic faults, most remain test-blind: The model sees only the source code and cannot reason about what existing tests already cover. We propose test-aware mutant generation, in which an LLM receives the problem statement, canonical solution and base tests in a single prompt, and must generate a nontrivial mutant that passes the base unit tests. We evaluate this approach across a set of five LLMs -- Gemini 3.1 Pro, Gemini 3 Flash, GPT 5.1 Codex Mini, GPT 4.1 Mini, Qwen3-32B -- on the HumanEval and MBPP benchmarks. The extended EvalPlus test suites serve as an automated oracle to verify whether surviving mutants represent genuine bugs. Test-aware prompting yields verified fault rates of 87.7% (HumanEval) and 79.1% (MBPP), meaning these mutants pass all base tests but are caught by the oracle. This vastly outperforms the matched test-blind prompting (which yields only 12.2% and 23.0%, respectively) and the traditional rule-based tool mutmut (4.4% and 5.7%). While fault subtlety (the fraction of extended tests a mutant fails) remains comparable across all three methods, test-awareness minimizes the computational cost per verified fault, compared to test-blind prompting. Exposing an LLM to existing unit tests shifts mutant generation from untargeted bug injection toward effective discovery of weaknesses in an existing test suite. Our work establishes a concrete foundation for future research to scale test-aware mutant generation to production-level environments.
Problem

Research questions and friction points this paper is trying to address.

mutation testing
test-aware mutant generation
large language models
test suite adequacy
test-blind
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Aware Mutant Generation
Large Language Models
Mutation Testing
Prompt Engineering
Test Suite Adequacy
N
Nils Kiele
Department of Electrical and Software Engineering, University of Calgary, AB, Canada
Z
Zainab Saad
Department of Electrical and Software Engineering, University of Calgary, AB, Canada
Z
Zirui Wang
Department of Electrical and Software Engineering, University of Calgary, AB, Canada
Steve Drew
Steve Drew
Assistant Professor at University of Calgary
Edge AIIoTMachine Learning
Samira Ebrahimi Kahou
Samira Ebrahimi Kahou
Associate Professor, University of Calgary/Mila/Canada CIFAR AI Chair
Machine LearningComputer VisionDeep LearningMultimodal LearningReinforcement Learning