test-driven generation

Designs and implements automated generation pipelines that produce artifacts (for example, programs or configurations), execute those artifacts with execution-based tests, and analyze test outcomes. Builds test harnesses, automated test generators and iteration control that regenerate or refine outputs until tests pass and surface diagnostic failure messages to guide further generation or developer fixes.

test-drivengeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.68
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$197K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

HarnessLLM: Automatic Testing Harness Generation via Reinforcement Learning

Nov 02, 2025
YL
Yujian Liu
🏛️ UC Santa Barbara | MIT-IBM Watson AI Lab | MIT CSAIL

Existing LLM-based automated test generation primarily produces static input-output assertion pairs, resulting in limited test diversity and insufficient debugging information. This work proposes a novel paradigm for generating executable test harnesses—supporting dynamic input construction and flexible output validation (e.g., invariant checking). Methodologically, we design a two-stage training framework: first, supervised fine-tuning (SFT) to teach LLMs the structural conventions of test scripts; second, reinforcement learning with a custom reward function (RLVR) to optimize test quality along dimensions such as correctness, coverage, and verifiability. Empirical evaluation demonstrates substantial improvements in defect detection rate and test strategy diversity; moreover, the generated harnesses support runtime extension to further enhance code generation fidelity. Our core contribution is the first systematic advancement of LLM-driven test generation—from static assertion pairs to fully executable, formally verifiable, and extensible test programs.

Creating harness code that synthesizes inputs and validates outputsGenerating diverse test cases beyond simple input-output pairsProviding comprehensive debugging information for program validation

Automatically Generating Single-Responsibility Unit Tests

Apr 08, 2025
GG
Geraldine Galindo-Gutierrez
🏛️ Universidad Católica Boliviana

Automated unit test generation often suffers from structural disorganization, poor comprehensibility, and low developer acceptance—particularly due to ambiguous relationships between test logic and assertions. To address this, we propose the first systematic integration of the Single Responsibility Principle (SRP) into search-based test generation, introducing an SRP-guided preprocessing step that structurally decouples coverage-target optimization from semantic clarity modeling. Our approach employs SRP-aware structural mutation operators and is rigorously evaluated via quantitative metrics (line/branch coverage, fault detection rate) and a developer empirical study. Results demonstrate statistically significant improvements in test comprehensibility (p < 0.01), with no degradation in coverage or fault detection performance, and a 37% increase in developer satisfaction. The core contribution is a novel paradigm for structured test generation that jointly ensures effectiveness and human-centered acceptability.

Enhancing test quality by applying single-responsibility principleEvaluating impact of structured tests on coverage and fault detectionImproving test structure for better understandability in generated tests

Execution-Feedback Driven Test Generation from SWE Issues

Aug 08, 2025
TA
Toufique Ahmed
🏛️ IBM Research

When target code is missing or erroneous, generating reproducible test cases becomes challenging due to the absence of a correct oracle. To address this, this paper proposes an execution-feedback-driven test generation method that, without relying on a correct implementation, dynamically captures runtime behavioral deviations and guides test inputs toward conditions triggering SWE (Software Engineering) issues via repair-oriented constraint solving. Implemented in the custom tool e-Otter++, the approach overcomes the traditional limitation of requiring correct-code execution feedback. Evaluated on the TDD-Bench Verified benchmark, it achieves an average failure-to-pass (F2P) rate of 63%, significantly outperforming state-of-the-art techniques. Its core contribution is the first construction of a closed-loop execution feedback mechanism specifically designed for scenarios involving erroneous or missing code—enabling high-precision, robust reproduction of SWE issues through automatically generated test cases.

Generating reproduction tests for SWE issues automaticallyLeveraging execution feedback without correct codeOvercoming missing or incorrect code in test generation

Intent Preserving Generation of Diverse and Idiomatic (Code-)Artifacts

Aug 05, 2025
OW
Oliver Westphal
🏛️ Universität Duisburg-Essen

To address the challenges of simultaneously generating semantically consistent yet stylistically diverse multi-artifact programming exercises—namely source code, test specifications, and natural language descriptions—this paper proposes a compositional generation framework grounded in abstract syntax building blocks. The framework defines reusable syntactic abstractions and integrates templated mapping with multi-objective instantiation to ensure intent preservation and cross-modal co-generation. Its key innovations include: (i) enabling style-controllable, diverse outputs while guaranteeing semantic consistency; and (ii) providing a highly configurable generation interface that substantially reduces customization effort for new tasks. Experimental evaluation demonstrates that the approach outperforms existing baselines across three critical dimensions: generation quality, output diversity, and system extensibility.

Automated generation of diverse, idiomatic code for programming exercisesCreating adaptable, non-monolithic generators from abstract building blocksManaging multiple related artifacts like specifications and descriptions

AI-powered test automation tools: A systematic review and empirical evaluation

Aug 31, 2024
VG
Vahid Garousi
🏛️ Queen's University Belfast | Testinium A. Ş. | ProSys MMC

Despite growing adoption of AI-driven test automation tools, their real-world efficacy—particularly in improving test efficiency, reducing maintenance costs, and enhancing defect detection—remains inadequately evaluated. Method: We conduct a systematic literature review identifying 55 tools and propose the first taxonomy of AI testing capabilities; further, we perform a dual-tool, dual-system empirical study on open-source projects, evaluating core functionalities including UI self-healing, visual testing, and intelligent test case generation. Contribution/Results: AI tools improve execution efficiency and reduce maintenance effort by over 30%, yet suffer from high false-positive rates, insufficient domain knowledge integration, and strong model dependency. This work establishes the first benchmarking framework for AI-based testing that jointly integrates a comprehensive capability taxonomy with multi-dimensional empirical validation—providing foundational guidance for developing robust, interpretable, and production-ready AI testing tools.

Compares AI tools with traditional methodsEvaluates AI-powered test automation toolsIdentifies AI features and limitations

Latest Papers

What's happening recently
View more

This study addresses a critical yet previously underexplored issue in large language model (LLM)-driven software development: the contamination of automatically generated tests by erroneous code. The authors systematically uncover and empirically validate this error propagation phenomenon, demonstrating that when tests are generated based on incorrect code within multi-step agent workflows—across diverse programming tasks and various prompting strategies, including chain-of-thought—the resulting tests exhibit significantly lower defect detection rates (14%) compared to independently generated tests (25%). These findings challenge the prevailing assumption that LLM-generated tests can serve as reliable, independent oracles, thereby highlighting the substantial risk of test bias in LLM-augmented development pipelines.

automated testing reliabilitycode-test consistencyerror propagation

This study addresses the unclear relationships among information sources, generation strategies, and quality evidence in test case generation using large language models (LLMs). Through a systematic literature review of 95 studies, this work constructs a multidimensional taxonomy and a benchmark analysis framework. Specifically, it proposes a four-dimensional classification system that elucidates how execution feedback influences oracle independence. Furthermore, it establishes a unified theoretical framework connecting the generation process with quality assessment. By identifying independent oracle evaluation as a critical yet underexplored dimension, this research formulates a future agenda centered on rigorous, oracle-independent quality measurement for LLM-generated test cases.

Large Language ModelsSoftware TestingTest Generation

This work addresses the challenge developers face in efficiently authoring CI/CD configurations due to limited DevOps expertise by proposing a large language model (LLM)-based, context-aware generation approach. The method leverages both natural language descriptions and repository structure to automatically produce accurate and executable pipeline configurations for platforms such as GitHub Actions and GitLab CI/CD. Integrated with automated validation and human-in-the-loop feedback mechanisms, this framework is the first to combine repository context understanding with natural language-driven configuration synthesis. Experimental results demonstrate that the approach significantly lowers the barrier to DevOps adoption, markedly improves the accuracy and validity of generated configurations, and substantially reduces manual configuration effort.

CI/CD pipeline configurationconfiguration errorsdeveloper productivity

Current large language model (LLM) agents struggle to precisely localize harness defects responsible for unreliable behaviors within failed execution trajectories, leading to broad and inefficient remediation strategies. This work proposes HarnessFix, a novel framework that enables the first precise diagnosis and structured repair of harness defects based on execution traces. By constructing a harness-aware trajectory intermediate representation (HTIR), HarnessFix integrates step-level provenance tracking, control-flow analysis, and defect aggregation to fine-grainedly attribute faulty behaviors to specific steps and harness components, subsequently generating specification-guided repair patches. Experimental results demonstrate that HarnessFix achieves performance gains of 15.2%–50.0% across four benchmarks, including SWE-Bench Verified, significantly outperforming both handcrafted and self-evolution baselines, while uncovering recurrent harness defect patterns in the ETCLOVG architecture.

execution tracesfailure diagnosisharness flaws

This work addresses the challenge that large language model (LLM) agents struggle to adapt at test time to distribution shifts, novel failure modes, or new tool interactions due to their execution pipelines being fixed prior to deployment. To overcome this limitation, the authors propose an unsupervised test-time evolution method that reframes adaptation as an optimization problem over executable control programs. By analyzing execution traces, the approach leverages population-based program evolution combined with an unsupervised proposer–discriminator mechanism to dynamically refine the control logic of ReAct-style agents—without updating model weights or relying on labeled data. Relying solely on frozen LLMs engaged in multi-role collaboration, the method achieves continuous, interpretable performance gains and significantly outperforms fixed-pipeline baselines on text-to-SQL, programming competition, and software engineering tasks.

execution tracesharness optimizationLLM agent

Hot Scholars

ZZ

Zibin Zheng

IEEE Fellow, Highly Cited Researcher, Sun Yat-sen University, China
BlockchainSmart ContractServices ComputingSoftware Reliability
JM

Jie M. Zhang

Lecturer (Assistant Professor), King's College London
LLMsSE4MLmachine learning testingmutation testing
ML

Mingwei Liu

Rutgers University
China laborhigh performance work systems
YL

Yuyu Luo

Assistant Professor, HKUST(GZ) / HKUST
Data AgentsLLM AgentsDatabaseText-to-SQL
QZ

Quanjun Zhang

Nanjing University of Science and Technology
Software EngineeringSoftware TestingAutomated Program Repair