Score
Designs and implements automated generation pipelines that produce artifacts (for example, programs or configurations), execute those artifacts with execution-based tests, and analyze test outcomes. Builds test harnesses, automated test generators and iteration control that regenerate or refine outputs until tests pass and surface diagnostic failure messages to guide further generation or developer fixes.
本文探讨了大型语言模型在软件工程中基于测试的方法,通过分析87篇研究文献,区分并比较了不同测试驱动任务的特点和机制,提出了未来研究方向。
Existing LLM-based automated test generation primarily produces static input-output assertion pairs, resulting in limited test diversity and insufficient debugging information. This work proposes a novel paradigm for generating executable test harnesses—supporting dynamic input construction and flexible output validation (e.g., invariant checking). Methodologically, we design a two-stage training framework: first, supervised fine-tuning (SFT) to teach LLMs the structural conventions of test scripts; second, reinforcement learning with a custom reward function (RLVR) to optimize test quality along dimensions such as correctness, coverage, and verifiability. Empirical evaluation demonstrates substantial improvements in defect detection rate and test strategy diversity; moreover, the generated harnesses support runtime extension to further enhance code generation fidelity. Our core contribution is the first systematic advancement of LLM-driven test generation—from static assertion pairs to fully executable, formally verifiable, and extensible test programs.
Automated unit test generation often suffers from structural disorganization, poor comprehensibility, and low developer acceptance—particularly due to ambiguous relationships between test logic and assertions. To address this, we propose the first systematic integration of the Single Responsibility Principle (SRP) into search-based test generation, introducing an SRP-guided preprocessing step that structurally decouples coverage-target optimization from semantic clarity modeling. Our approach employs SRP-aware structural mutation operators and is rigorously evaluated via quantitative metrics (line/branch coverage, fault detection rate) and a developer empirical study. Results demonstrate statistically significant improvements in test comprehensibility (p < 0.01), with no degradation in coverage or fault detection performance, and a 37% increase in developer satisfaction. The core contribution is a novel paradigm for structured test generation that jointly ensures effectiveness and human-centered acceptability.
When target code is missing or erroneous, generating reproducible test cases becomes challenging due to the absence of a correct oracle. To address this, this paper proposes an execution-feedback-driven test generation method that, without relying on a correct implementation, dynamically captures runtime behavioral deviations and guides test inputs toward conditions triggering SWE (Software Engineering) issues via repair-oriented constraint solving. Implemented in the custom tool e-Otter++, the approach overcomes the traditional limitation of requiring correct-code execution feedback. Evaluated on the TDD-Bench Verified benchmark, it achieves an average failure-to-pass (F2P) rate of 63%, significantly outperforming state-of-the-art techniques. Its core contribution is the first construction of a closed-loop execution feedback mechanism specifically designed for scenarios involving erroneous or missing code—enabling high-precision, robust reproduction of SWE issues through automatically generated test cases.
To address the challenges of simultaneously generating semantically consistent yet stylistically diverse multi-artifact programming exercises—namely source code, test specifications, and natural language descriptions—this paper proposes a compositional generation framework grounded in abstract syntax building blocks. The framework defines reusable syntactic abstractions and integrates templated mapping with multi-objective instantiation to ensure intent preservation and cross-modal co-generation. Its key innovations include: (i) enabling style-controllable, diverse outputs while guaranteeing semantic consistency; and (ii) providing a highly configurable generation interface that substantially reduces customization effort for new tasks. Experimental evaluation demonstrates that the approach outperforms existing baselines across three critical dimensions: generation quality, output diversity, and system extensibility.
Despite growing adoption of AI-driven test automation tools, their real-world efficacy—particularly in improving test efficiency, reducing maintenance costs, and enhancing defect detection—remains inadequately evaluated. Method: We conduct a systematic literature review identifying 55 tools and propose the first taxonomy of AI testing capabilities; further, we perform a dual-tool, dual-system empirical study on open-source projects, evaluating core functionalities including UI self-healing, visual testing, and intelligent test case generation. Contribution/Results: AI tools improve execution efficiency and reduce maintenance effort by over 30%, yet suffer from high false-positive rates, insufficient domain knowledge integration, and strong model dependency. This work establishes the first benchmarking framework for AI-based testing that jointly integrates a comprehensive capability taxonomy with multi-dimensional empirical validation—providing foundational guidance for developing robust, interpretable, and production-ready AI testing tools.
This study addresses a critical yet previously underexplored issue in large language model (LLM)-driven software development: the contamination of automatically generated tests by erroneous code. The authors systematically uncover and empirically validate this error propagation phenomenon, demonstrating that when tests are generated based on incorrect code within multi-step agent workflows—across diverse programming tasks and various prompting strategies, including chain-of-thought—the resulting tests exhibit significantly lower defect detection rates (14%) compared to independently generated tests (25%). These findings challenge the prevailing assumption that LLM-generated tests can serve as reliable, independent oracles, thereby highlighting the substantial risk of test bias in LLM-augmented development pipelines.
This study addresses the unclear relationships among information sources, generation strategies, and quality evidence in test case generation using large language models (LLMs). Through a systematic literature review of 95 studies, this work constructs a multidimensional taxonomy and a benchmark analysis framework. Specifically, it proposes a four-dimensional classification system that elucidates how execution feedback influences oracle independence. Furthermore, it establishes a unified theoretical framework connecting the generation process with quality assessment. By identifying independent oracle evaluation as a critical yet underexplored dimension, this research formulates a future agenda centered on rigorous, oracle-independent quality measurement for LLM-generated test cases.
This work addresses the challenge developers face in efficiently authoring CI/CD configurations due to limited DevOps expertise by proposing a large language model (LLM)-based, context-aware generation approach. The method leverages both natural language descriptions and repository structure to automatically produce accurate and executable pipeline configurations for platforms such as GitHub Actions and GitLab CI/CD. Integrated with automated validation and human-in-the-loop feedback mechanisms, this framework is the first to combine repository context understanding with natural language-driven configuration synthesis. Experimental results demonstrate that the approach significantly lowers the barrier to DevOps adoption, markedly improves the accuracy and validity of generated configurations, and substantially reduces manual configuration effort.
Current large language model (LLM) agents struggle to precisely localize harness defects responsible for unreliable behaviors within failed execution trajectories, leading to broad and inefficient remediation strategies. This work proposes HarnessFix, a novel framework that enables the first precise diagnosis and structured repair of harness defects based on execution traces. By constructing a harness-aware trajectory intermediate representation (HTIR), HarnessFix integrates step-level provenance tracking, control-flow analysis, and defect aggregation to fine-grainedly attribute faulty behaviors to specific steps and harness components, subsequently generating specification-guided repair patches. Experimental results demonstrate that HarnessFix achieves performance gains of 15.2%–50.0% across four benchmarks, including SWE-Bench Verified, significantly outperforming both handcrafted and self-evolution baselines, while uncovering recurrent harness defect patterns in the ETCLOVG architecture.
This work addresses the challenge that large language model (LLM) agents struggle to adapt at test time to distribution shifts, novel failure modes, or new tool interactions due to their execution pipelines being fixed prior to deployment. To overcome this limitation, the authors propose an unsupervised test-time evolution method that reframes adaptation as an optimization problem over executable control programs. By analyzing execution traces, the approach leverages population-based program evolution combined with an unsupervised proposer–discriminator mechanism to dynamically refine the control logic of ReAct-style agents—without updating model weights or relying on labeled data. Relying solely on frozen LLMs engaged in multi-role collaboration, the method achieves continuous, interpretable performance gains and significantly outperforms fixed-pipeline baselines on text-to-SQL, programming competition, and software engineering tasks.