Score
Designs and implements test suites, automated harnesses, and analysis procedures that detect interface and behavioral regressions between versions of software artifacts—covering APIs, ML models, and similar components—by exercising inputs, checking invariants, verifying operator correctness under aggregation, and running integration and regression scenarios to assess backward/forward compatibility.
This study addresses the imbalance in the test pyramid—characterized by an overreliance on coarse-grained integration and system tests, which leads to difficulties in fault localization and slow execution—by proposing, for the first time, a method to automatically generate unit tests from existing integration tests. The approach combines static and dynamic analysis to automatically isolate component dependencies and enhance coverage at the unit level. Implemented as a Node.js tool and evaluated on twelve open-source JavaScript projects, the technique produces high-quality unit tests that significantly improve test suite structure, thereby increasing both testing efficiency and maintainability.
This study addresses the challenges of regression testing in remote and hybrid work environments, where communication, coordination, and quality assurance are increasingly complex. Through qualitative interviews with 20 software practitioners, complemented by process analysis, tool integration assessment, and coding of collaborative practices, the research systematically investigates the sociotechnical evolution of regression testing in distributed settings. Findings indicate that while core testing phases remain largely stable, teams increasingly rely on documentation, automation, and integrated toolchains to sustain effectiveness. Standardized reporting formats, shared repositories, and traceability mechanisms significantly mitigate collaboration barriers inherent in remote work. The study offers novel insights and practical guidance for ensuring software quality in geographically dispersed development contexts.
This work addresses the limitations of traditional structural coverage metrics in embedded software testing, which are often confined to the unit level and fail to reflect true coverage completeness in integration and system testing. Instrumentation-based approaches risk perturbing runtime behavior, while pure tracing techniques suffer from unreliability under high compiler optimization. To overcome these challenges, the paper proposes an integration-test-driven coverage strategy featuring a novel “integration-first” closed-loop workflow. By synergistically combining embedded tracing with hybrid runtime analysis (hRA) to preserve semantic boundaries, and leveraging source-to-target mapping for evidential traceability alongside Hyper Coverage for cross-variant merging, the approach establishes a unified evidence-integration mechanism. Evaluated on -O3-optimized release binaries, it reliably achieves branch, condition, and MC/DC coverage measurements and precisely identifies source code lines consistently uncovered across all variants, thereby significantly enhancing confidence in the test completeness of embedded systems.
Traditional regression testing theory fails in agile and continuous integration settings due to the dynamic, time-ordered nature of continuous builds. Method: This paper proposes a formal modeling framework based on time-ordered build chains, representing continuous build sequences as temporally constrained build-tuple chains and formally defining the novel concept of the “regression testing window”—unifying classical two-version and multi-version continuous testing scenarios. The framework enables efficient verification of correctness and completeness of regression testing within bounded time and supports rigorous formal verification via logical deduction. Contribution/Results: Experimental evaluation demonstrates that the model successfully characterizes and verifies two state-of-the-art agile regression testing algorithms. It exhibits strong expressive power, requires no auxiliary assumptions, ensures theoretical soundness, and is directly deployable in practice.
Empirical software engineering lacks standardized tooling to support Test-Driven Software Experiments (TDSE)—i.e., executing subject software and observing/analyzing its runtime behavior. To address this gap, we propose LASSO, the first general-purpose TDSE framework explicitly designed for runtime semantic assessment. LASSO integrates a domain-specific language (DSL), executable experiment scripts, dynamic behavioral instrumentation and analysis infrastructure, and an open-source web platform with interactive visualization. It enables self-contained, reusable, and extensible empirical studies on the reliability of LLM-generated code. An empirical evaluation conducted using LASSO validates its effectiveness in facilitating rigorous, reproducible experimentation. The platform has been publicly released as open-source software and adopted in practice, demonstrably improving experimental development efficiency, reproducibility, and interpretability.
This work addresses the limitation of conventional execution coverage in UI component testing, which fails to verify whether tests adequately capture behavioral relationships implied by APIs and documentation. The paper proposes the first evaluation framework based on inferred metamorphic relations (MRs): it automatically derives MRs using a UI-specific taxonomy from source code and documentation, aligns test executions to these MRs through deterministic and semantic analysis, and introduces relation-level MR coverage as a novel metric. By treating inferred MRs as empirical benchmarks for behavioral validation, the approach exposes verification gaps invisible to traditional coverage metrics—particularly in weak-oracle scenarios. Empirical results across three LLM configurations show MR coverage ranging only from 42.5% to 47.6%, substantially lower than MR reachability; uncovered MRs are predominantly of the weak-oracle type, demonstrating that MR coverage meaningfully complements conventional metrics and offers practical utility in fault detection and issue mapping.
This work addresses the challenge that large language model (LLM) agents struggle to adapt at test time to distribution shifts, novel failure modes, or new tool interactions due to their execution pipelines being fixed prior to deployment. To overcome this limitation, the authors propose an unsupervised test-time evolution method that reframes adaptation as an optimization problem over executable control programs. By analyzing execution traces, the approach leverages population-based program evolution combined with an unsupervised proposer–discriminator mechanism to dynamically refine the control logic of ReAct-style agents—without updating model weights or relying on labeled data. Relying solely on frozen LLMs engaged in multi-role collaboration, the method achieves continuous, interpretable performance gains and significantly outperforms fixed-pipeline baselines on text-to-SQL, programming competition, and software engineering tasks.
In enterprise microservice regression testing, QA engineers often lack up-to-date documentation and must rely on real user traffic to reconstruct business scenarios; however, transforming such traffic into replayable test cases with stable assertions is labor-intensive and error-prone. This work proposes NL2Test, a method that combines the semantic understanding of large language models with deterministic algorithms to generate executable API test cases end-to-end from natural language scenario descriptions and captured execution traces. NL2Test automatically slices request sequences, reconstructs data dependencies, masks non-deterministic fields, and produces business-aligned, reliable assertions. Evaluated on 51 industrial scenarios, it achieves an exact match rate of 82.4%, with 98.0% of generated test cases becoming functional after minor tuning. During a nine-month production deployment, it produced 3,196 test cases, of which 85.4% were accepted and integrated into the codebase.
Existing runtime harnesses for programming agents suffer from either oversimplification or excessive complexity, lacking a clear and concise architectural paradigm. This work proposes a harness design centered on the request lifecycle, explicitly delineating three core boundaries: model, execution, and state. By orchestrating a structured sequence—comprising context construction, model decision-making, environmental action, observation feedback, and state continuation—the design enables cross-request state persistence and continual self-improvement through bootstrapping. We implement this paradigm in Coderlet, an open-source prototype system, demonstrating its efficacy in coordinating code generation, environment interaction, and state management. The resulting framework provides a scalable foundation for building high-performance programming agents.
This work addresses the challenge of effectively triggering behavioral discrepancies caused by code changes, which often leads to missed regression defects in existing testing approaches. The authors propose a novel LLM-driven differential testing method tailored to code modifications, which innovatively integrates static call graph analysis, project documentation understanding, and large language model (LLM) generation capabilities. Guided by a joint coverage feedback mechanism, the approach iteratively refines test case generation to simultaneously enhance coverage of changed code across both old and new software versions. Evaluated on 463 pull requests, the method successfully exposed behavioral differences in 78.2% of them, achieved an average joint coverage of 90.7%, uncovered 99 additional issues compared to baseline techniques, improved coverage by 12.5–15.6 percentage points, and effectively detected regression defects overlooked by prior methods.