Score
Designs, implements, and operates automated systems, frameworks, and pipelines that execute and manage regression test suites, generate and run regression test cases, and perform automated validation and verification of functional and performance regressions. Builds tooling for regression monitoring, detection, tracking, reporting, alerting, and triage so teams can identify, reproduce, and verify fixes to regressions.
This study addresses the challenges of regression testing in remote and hybrid work environments, where communication, coordination, and quality assurance are increasingly complex. Through qualitative interviews with 20 software practitioners, complemented by process analysis, tool integration assessment, and coding of collaborative practices, the research systematically investigates the sociotechnical evolution of regression testing in distributed settings. Findings indicate that while core testing phases remain largely stable, teams increasingly rely on documentation, automation, and integrated toolchains to sustain effectiveness. Standardized reporting formats, shared repositories, and traceability mechanisms significantly mitigate collaboration barriers inherent in remote work. The study offers novel insights and practical guidance for ensuring software quality in geographically dispersed development contexts.
Traditional regression testing theory fails in agile and continuous integration settings due to the dynamic, time-ordered nature of continuous builds. Method: This paper proposes a formal modeling framework based on time-ordered build chains, representing continuous build sequences as temporally constrained build-tuple chains and formally defining the novel concept of the “regression testing window”—unifying classical two-version and multi-version continuous testing scenarios. The framework enables efficient verification of correctness and completeness of regression testing within bounded time and supports rigorous formal verification via logical deduction. Contribution/Results: Experimental evaluation demonstrates that the model successfully characterizes and verifies two state-of-the-art agile regression testing algorithms. It exhibits strong expressive power, requires no auxiliary assumptions, ensures theoretical soundness, and is directly deployable in practice.
Existing regression testing approaches for Robot Operating System–based Autonomous Systems (ROSAS) lack systematic optimization strategies to address challenges arising from dynamic behaviors, multimodal perception, asynchronous architectures, and real-time safety constraints. Method: This work proposes the first taxonomy for ROSAS regression test optimization, structured along three dimensions—test prioritization, minimization, and selection—and introduces novel techniques: frame-to-vector coverage metrics, multi-source foundation model–driven test generation, and neuro-symbolic reasoning–based verification. The framework is developed through a systematic literature review of 122 papers, taxonomy modeling, and technology roadmap design. Contribution/Results: The study delivers a scalable, reusable regression testing optimization framework; establishes the first theoretical taxonomy for ROSAS testing research; and provides industry with the inaugural practical guideline for ROSAS regression testing—bridging the gap between academic rigor and industrial applicability.
To address test redundancy, high feedback latency, and inconsistent pre- vs. post-commit test selection objectives in large-scale multilingual monorepos, this paper proposes the first pipeline-aware, bi-objective reinforcement learning framework for regression test optimization: failure detection is prioritized during pre-commit testing, while flaky-change identification is emphasized post-commit. The method operates entirely on language-agnostic features, integrating pipeline semantic modeling with online log analysis to support dynamically evolving industrial test suites. Evaluated on 20 weeks of real-world CI data, it achieves significantly reduced average feedback latency, a 32% improvement in pre-commit test selection precision, and a 41% reduction in false positives—without requiring expensive features such as code coverage.
To address high manual effort, delayed response, and incomplete coverage in industrial software test maintenance, this paper proposes and empirically validates two novel multi-agent architectures leveraging large language models (LLMs) to accurately identify test cases requiring maintenance—and to assist in their automated addition, deletion, or modification—following code changes. The study systematically distills critical triggering conditions and practical deployment constraints for LLMs in industrial settings, and conducts empirical evaluation using real-world data from Ericsson AB. Results demonstrate significant reductions in manual test maintenance effort, improved timeliness of issue response, and enhanced test coverage completeness. This work represents the first application of multi-agent collaboration to industrial-scale test maintenance, establishing a reusable methodological framework and empirical foundation for deploying LLMs in high-reliability software engineering contexts.
In enterprise microservice regression testing, QA engineers often lack up-to-date documentation and must rely on real user traffic to reconstruct business scenarios; however, transforming such traffic into replayable test cases with stable assertions is labor-intensive and error-prone. This work proposes NL2Test, a method that combines the semantic understanding of large language models with deterministic algorithms to generate executable API test cases end-to-end from natural language scenario descriptions and captured execution traces. NL2Test automatically slices request sequences, reconstructs data dependencies, masks non-deterministic fields, and produces business-aligned, reliable assertions. Evaluated on 51 industrial scenarios, it achieves an exact match rate of 82.4%, with 98.0% of generated test cases becoming functional after minor tuning. During a nine-month production deployment, it produced 3,196 test cases, of which 85.4% were accepted and integrated into the codebase.
本文探讨了大型语言模型在软件工程中基于测试的方法,通过分析87篇研究文献,区分并比较了不同测试驱动任务的特点和机制,提出了未来研究方向。
This study investigates the risk that regression tests generated by large language models (LLMs) may inadvertently encode software defects, allowing erroneous behaviors to evade detection. By integrating LLMs, automated test generation, and software repository mining techniques, this work systematically evaluates the likelihood of LLM-generated tests encoding incorrect behaviors during real-world software evolution. It reveals how such tests induce a “fault-forcing” phenomenon, thereby introducing a novel form of technical debt. Empirical findings demonstrate that 8%–17% of generated tests enforce fault preservation, with most remaining latent in codebases over extended periods and only a negligible fraction detected through manual testing. To our knowledge, this research provides the first quantitative assessment of this risk, offering critical empirical evidence to guide and standardize LLM-assisted testing practices.
This work addresses the challenge of efficiently deriving the most general preconditions required to achieve a goal in automated planning with axioms, a task where traditional logical regression suffers from high computational complexity. To overcome this limitation, the paper proposes an approximate logical regression method that restricts preconditions to partial states, thereby avoiding redundant axiom computations while efficiently generating minimal partial states. This approach constitutes the first efficient approximation of logical regression in axiom-rich planning domains, significantly enhancing the generalization capability of partial states and the robustness of execution monitoring. Empirical results across multiple planning domains demonstrate up to a 70% reduction in the number of monitored variables and over 50% task recovery success under unexpected environmental perturbations.
This study addresses a critical yet overlooked issue in large language model (LLM) agents: while procedural skills improve average task success rates, they often induce “regression”—causing previously solvable tasks to fail. Through controlled experiments on nearly 6,000 office automation tasks, this work quantifies and disentangles the dual effects of skill integration, introducing the concept of a “regression tax.” The findings reveal that skill reliability hinges more on grounding and verification mechanisms than on procedural logic itself; the superiority of optimal skills stems primarily from their lower regression rates; and most regression failures can be mitigated through enhanced verification. These insights establish a new paradigm for skill design, grounded in empirical evidence and emphasizing robustness over mere capability expansion.