On the Risks of using LLM-Generated Tests for Regression Testing

πŸ“… 2026-10-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates the risk that regression tests generated by large language models (LLMs) may inadvertently encode software defects, allowing erroneous behaviors to evade detection. By integrating LLMs, automated test generation, and software repository mining techniques, this work systematically evaluates the likelihood of LLM-generated tests encoding incorrect behaviors during real-world software evolution. It reveals how such tests induce a β€œfault-forcing” phenomenon, thereby introducing a novel form of technical debt. Empirical findings demonstrate that 8%–17% of generated tests enforce fault preservation, with most remaining latent in codebases over extended periods and only a negligible fraction detected through manual testing. To our knowledge, this research provides the first quantitative assessment of this risk, offering critical empirical evidence to guide and standardize LLM-assisted testing practices.
πŸ“ Abstract
Software is under constant evolution: developers continuously add features, fix bugs, and refactor code, and any of these changes may break existing functionality. Regression testing guards against such effects by capturing expected behavior in test cases. LLM-based test generation aims to automate this process by generating regression tests directly from the code under test. This is beneficial when the implementation is correct, but problematic when the code contains faults: the generated tests may then encode and preserve incorrect behavior. To investigate this risk, we apply LLM-based regression test generation to pull requests merged into the main branch of software projects and study the impact of the generated tests on subsequent project evolution. We distinguish between fault-revealing tests, which assert correctly implemented behavior, and fault-enforcing tests, which assert faulty behavior. Across 145 pull requests from SciPy, Qiskit, and pandas, 8%-17% of the generated tests are fault-enforcing, while only 2.4%-4.8% reveal faults. Fault-enforcing tests persist over time: after several subsequent commits, 83%-91% of them are still relevant and pass. They also accumulate: when the faults of all pull requests are combined in one codebase, 83%-92% remain enforced at the end of the commit history, and the developer-written test suite detects only 14%-30% of them. Our results reveal a fundamental risk of LLM-generated regression tests: without manual validation, they may encode faulty behavior as expected behavior, allowing bugs to persist across software revisions and largely evade developer-maintained test suites. LLM-based regression testing can thus give rise to a new form of technical debt.
Problem

Research questions and friction points this paper is trying to address.

LLM-generated tests
regression testing
fault-enforcing tests
software evolution
technical debt
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-generated tests
Regression testing
Fault-enforcing tests
Technical debt
Software evolution
πŸ’Ό Related Jobs
No related jobs found.
M
Mohammadali Charoosaei
SnT, University of Luxembourg, Luxembourg
C
Cedric Richter
SnT, University of Luxembourg, Luxembourg
Mike Papadakis
Mike Papadakis
Associate professor, University of Luxembourg
Software EngineeringMutation TestingSoftware TestingSoftware EvolutionSBSE