Score
Designs and produces formal test plans and the associated execution procedures that specify scope, objectives, entry/exit and pass/fail criteria, test cases and test data, environment and setup, resources and schedule, and traceability to requirements. Documents step‑by‑step execution instructions, logging, reporting and defect-handling processes so tests can be carried out, evaluated, and repeated consistently.
本文探讨了大型语言模型在软件工程中基于测试的方法,通过分析87篇研究文献,区分并比较了不同测试驱动任务的特点和机制,提出了未来研究方向。
Existing BPMN+DMN process models lack semantic-level automated verification; mainstream tools support only syntactic validation, while behavioral errors require manual execution and debugging, and model transformations remain opaque. Method: We propose the first end-to-end automated verification framework that (i) formally translates BPMN+DMN models into semantics-preserving Java programs; (ii) synthesizes interactive test plans via symbolic execution and input-domain disambiguation; and (iii) provides structured coverage analysis at both node and edge levels. Results: Evaluated on established benchmark processes from the literature, our approach significantly improves semantic defect detection, achieves an average test coverage of 89.3%, and accelerates verification by over 20× compared to manual methods.
Real-time testing in cloud service production environments risks interfering with live traffic, compromising service reliability and SLA compliance. Method: This paper proposes an automated test planning framework that jointly models test configuration selection, deployment planning, and execution scheduling as a risk-constrained optimization problem. It integrates lightweight online traffic forecasting with dynamic risk mitigation strategies to coordinate test execution with production workload fluctuations. Contribution/Results: Compared to manual approaches, the method significantly reduces configuration errors and service disruption risks. In real-world cloud service deployments, it decreases SLA violations induced by testing by 62% and shortens average test preparation time by 57%, while maintaining full test coverage and effectiveness. The framework provides a scalable, resource-aware automation solution for safe and efficient online testing of large-scale distributed systems.
Existing search-based software testing (SBST) and large language model (LLM)-based approaches struggle to cover hard-to-test branches involving complex object construction and cross-procedural dependencies. To address this, we propose a semantics-aware LLM-based test generation method that integrates static program analysis with feedback-driven prompt engineering. Our approach introduces a novel three-stage mechanism: (1) extraction of realistic invocation contexts, (2) identification of cross-procedural dependencies, and (3) counterexample-guided iterative refinement—enabling LLMs to precisely comprehend target method semantics and constraints. By jointly modeling control and data flows, incorporating counterexample feedback for learning, and dynamically reconstructing prompts, our method achieves targeted triggering of hard-to-cover branches. Evaluated on 27 open-source Python projects, it improves average branch coverage by 31.39% over SBST and by 22.22% over LLM baselines, significantly enhancing testability of complex execution paths.
To address the challenges of automated test execution across projects, programming languages, build systems, and testing frameworks, this paper proposes ExecutionAgent—a large language model (LLM)-based autonomous agent. It employs a meta-prompt-driven system interaction paradigm to parse source code, perform environment-aware configuration, and autonomously generate test scripts for arbitrary open-source projects, supporting feedback-guided iterative debugging. Its novel zero-shot adaptation mechanism requires no predefined rules or human intervention, ensuring compatibility with 14 programming languages and mainstream ecosystems. Evaluated on 50 heterogeneous projects, ExecutionAgent successfully executed 33 test suites with only 7.5% result deviation from ground truth, achieving 6.6× higher performance than state-of-the-art methods. The average per-project execution time is 74 minutes, with an LLM inference cost of just $0.16.
This work addresses the high cost of manually writing formal specifications and the limitations of existing large language model (LLM)-based approaches that require white-box access to source code, thereby posing intellectual property and deployment constraints. The authors propose a black-box-driven method that leverages only test code and dynamic execution traces to generate candidate Java Modeling Language (JML) specifications via an LLM. These candidates are locally validated using bounded model checking, and an iterative feedback loop refines them based on verification outcomes. This approach is the first to enable fully automated formal specification generation without any access to the program’s internal structure. Evaluated on the SpecGenBench benchmark, it demonstrates that test-derived information effectively guides specification synthesis, while also highlighting critical challenges in checker compatibility and diagnostic feedback, substantially enhancing industrial applicability.
本文提出记录和利用执行轨迹来持久化和分析操作的内部行为,解决VDM-SL规范中操作内部行为观察问题。
This study addresses the unclear relationships among information sources, generation strategies, and quality evidence in test case generation using large language models (LLMs). Through a systematic literature review of 95 studies, this work constructs a multidimensional taxonomy and a benchmark analysis framework. Specifically, it proposes a four-dimensional classification system that elucidates how execution feedback influences oracle independence. Furthermore, it establishes a unified theoretical framework connecting the generation process with quality assessment. By identifying independent oracle evaluation as a critical yet underexplored dimension, this research formulates a future agenda centered on rigorous, oracle-independent quality measurement for LLM-generated test cases.
This work addresses the challenge of ensuring program correctness in natural language-to-code generation, which is often hindered by the absence of high-quality formal specifications. The authors propose VeriSpecGen, a framework that decomposes natural language requirements into atomic clauses through a traceable refinement mechanism, generates requirement-driven tests with explicit traceability mappings, and synthesizes formal specifications aligned with user intent by localizing and repairing faulty clauses upon verification failure. Integrating large language models (e.g., Claude Opus 4.5) with the Lean proof assistant, the approach leverages refinement trajectories to generate 343K training samples, substantially enhancing model generalization and reasoning capabilities. Evaluated on the VERINA SpecGen benchmark, VeriSpecGen achieves an accuracy of 86.6%, outperforming the best baseline by up to 31.8 percentage points and demonstrating a relative improvement of 62–106% in specification synthesis performance.
This study addresses the lack of traceable, structured linkage between high-level requirements and low-level automated testing in AI-enabled cyber-physical systems, which hinders compliance with regulatory demands for verifiable evidence. To bridge this gap, the paper introduces VNVSpec, a novel framework that enables end-to-end automated traceability and closed-loop verification from high-level engineering requirements to test cases. VNVSpec employs machine-readable verification and validation (V&V) specifications to support requirement ingestion, quality checks, metric-driven decomposition, test result association, and generation of audit-ready reports, all integrated into a continuous integration pipeline. Empirical evaluation demonstrates that the approach verifies 36 requirements against 449 tests in linear time, scales to tens of thousands of artifacts, and is fully reproducible through open-sourced code, test suites, and benchmark scripts.