Score
Designs and implements testing frameworks and test suites that execute multiple implementations or translations of the same program on the same inputs and automatically compare their outputs to detect semantic divergences, miscompilations, and regressions. Work includes generating or selecting inputs/kernels, running cross-compiler or cross-implementation comparisons, triaging differences, and isolating minimal reproducible cases that expose incorrect or divergent behavior.
This study addresses the imbalance in the test pyramid—characterized by an overreliance on coarse-grained integration and system tests, which leads to difficulties in fault localization and slow execution—by proposing, for the first time, a method to automatically generate unit tests from existing integration tests. The approach combines static and dynamic analysis to automatically isolate component dependencies and enhance coverage at the unit level. Implemented as a Node.js tool and evaluated on twelve open-source JavaScript projects, the technique produces high-quality unit tests that significantly improve test suite structure, thereby increasing both testing efficiency and maintainability.
Existing compiler testing techniques are often ill-suited for transpilers, as they typically lack multiple equivalent implementations and may produce non-executable output code. This work introduces metamorphic testing to transpiler validation by proposing the notion of “mutation consistency”: it defines metamorphic relations at the source-code level to verify whether structurally consistent and expected changes manifest in the generated code when the input DSL program undergoes semantics-preserving mutations. This approach enables defect detection without requiring execution of the generated code. We develop a mutation-based modeling method for metamorphic relations, a source-level structural consistency analysis mechanism, and implement an automated tool, MCP-Tester. Evaluated on real-world technology migration cases, our method effectively uncovers transpiler bugs that elude conventional fuzzing approaches.
This work addresses the problem of detecting functional differences between two program versions by automatically generating difference-exposing test cases (DETs). We propose an execution-feedback-driven iterative generation method leveraging large language models (LLMs), introducing a novel mechanism that dynamically incorporates runtime execution outcomes into prompt engineering to enable closed-loop test-case optimization. Our approach integrates differential testing, Python program analysis, and dynamic feedback–augmented prompting. Evaluated on 1,535 program pairs from Codeforces, our method achieves an 81.7% DET generation rate—substantially outperforming Pynguin (4.9%) and Differential Prompting (37.3%). These results empirically validate the effectiveness and generalizability of execution-driven LLM prompting for program behavior discrepancy detection.
To address the “coverage plateau” phenomenon—where large language models (LLMs) stall in regression test generation due to lack of program execution awareness—this paper proposes TestWeaver. Methodologically, TestWeaver establishes an execution-aware test generation paradigm: (1) backward slicing suppresses model hallucination by constraining generation to relevant program slices; (2) dynamically retrieved neighbor test cases with similar control-flow graphs provide contextual execution information; and (3) inline annotations inject critical variable states to enhance LLM comprehension of program behavior. The framework integrates lightweight program analysis, control-flow modeling, and LLM inference—requiring no model retraining or fine-tuning. Empirical evaluation demonstrates that TestWeaver significantly accelerates coverage convergence and outperforms existing LLM-based baselines in both fault detection rate and path coverage depth.
This work addresses the challenge of detecting cross-language compilation bugs in multi-language JVM applications, which arise from semantic discrepancies between languages and are largely overlooked by existing compiler testing approaches focused on single-language settings. To bridge this gap, the authors propose the first differential testing framework tailored for cross-language JVM compilation. The approach leverages Kotlin’s unified intermediate representation (IR) to synthesize cross-language test programs and introduces seven custom mutation operators to enhance test diversity. This methodology enables the first systematic differential testing of multi-language JVM compilation scenarios, uncovering 32 confirmed bugs across five major compilers—Kotlin, Groovy, Scala 2/3, and Java. The findings demonstrate that the framework effectively exposes and mitigates semantic inconsistencies at language boundaries, significantly improving compiler reliability.
This study investigates whether large language models (LLMs) genuinely comprehend program semantics when generating unit tests or merely rely on superficial syntactic patterns. To this end, the authors propose the first systematic evaluation of LLMs’ regression-awareness in software evolution scenarios, introducing an automated mutation-driven framework that distinguishes semantic-altering changes (SAC) from semantic-preserving changes (SPC). An empirical analysis across 22,374 program variants involving eight prominent LLMs reveals that while these models achieve up to 79% line coverage on original programs, their performance degrades significantly after code evolution. Notably, over 99% of failing tests generated for evolved programs still pass on the original versions, indicating that LLMs are highly sensitive to syntactic modifications yet lack deep semantic understanding. This work provides critical empirical evidence regarding the reliability and limitations of LLM-based test generation.
为确保MongoDB在多种编程语言中的客户端库行为一致,开发了基于YAML的统一测试格式,减少了非一致性错误。
This study addresses the unclear relationships among information sources, generation strategies, and quality evidence in test case generation using large language models (LLMs). Through a systematic literature review of 95 studies, this work constructs a multidimensional taxonomy and a benchmark analysis framework. Specifically, it proposes a four-dimensional classification system that elucidates how execution feedback influences oracle independence. Furthermore, it establishes a unified theoretical framework connecting the generation process with quality assessment. By identifying independent oracle evaluation as a critical yet underexplored dimension, this research formulates a future agenda centered on rigorous, oracle-independent quality measurement for LLM-generated test cases.
This study addresses the challenges of high cost, error-proneness, and defect propagation in cross-repository code and test reuse during software refactoring. Through action research, the authors conduct bidirectional empirical analyses on real-world cases such as Soot/SootUp and FindBugs/SpotBugs, identifying for the first time the bidirectional reuse requirements and semantic reuse patterns inherent in refactoring scenarios. They propose a semantic alignment–based code mapping approach coupled with a hierarchical, extensible clone detection mechanism. Experimental results demonstrate that their method reduces irrelevant clones by 33%–99% on average and achieves a benchmark precision of 86%. The practical impact is further evidenced by five reported issues and ten pull requests submitted to open-source communities, eight of which have already been merged, confirming the approach’s effectiveness and applicability.
This work addresses the challenge of semantic inconsistencies between source code and its natural language documentation, which often lead to software defects and increased maintenance costs. To reduce false positives in inconsistency detection, the authors propose a large language model–based dual-validation mechanism that simultaneously generates unit tests and a reference implementation from the documentation. A semantic inconsistency is flagged only when the original code fails the generated tests while the reference implementation passes them. Integrating test generation, code synthesis, execution-based validation, and multi-language (Java/C#/Rust) static and dynamic analysis, the approach demonstrates high precision on a benchmark of 985 code-document pairs and uncovers 13 previously unknown inconsistencies in real-world open-source projects, ten of which have been confirmed and fixed by developers, significantly enhancing both accuracy and practical utility.