Score
Designs and builds test artifacts and infrastructure for software and data-transformations, including unit tests, test execution and automation frameworks, load and user tests, and automated test-generation pipelines; implements metamorphic testing approaches—defining metamorphic relations, synthesizing metamorphic tests, and producing oracle-independent test functions—to validate properties and detect semantic conflicts. Analyzes and improves test effectiveness by measuring coverage and sensitivity, validating property-preservation of reduced inputs, and optimizing test synthesis, execution, reduction, and performance workflows.
This work addresses the challenge of validating query-based systems, which often lack oracle outputs, and overcomes the inefficiency of traditional manual testing. The authors propose a metamorphic testing approach enhanced with a relational prompting mechanism that automatically derives metamorphic relations by uncovering semantic dependencies among inputs, eliminating the need for predefined test cases or ground-truth outputs. By introducing relational prompting into metamorphic testing for the first time, the method substantially reduces reliance on domain-specific knowledge and integrates effectively with complementary techniques such as combinatorial testing and fuzzing. Empirical evaluation on real-world web applications demonstrates that the proposed framework significantly enhances the automation, efficiency, and practical feasibility of defect detection in query-based systems.
Existing code coverage metrics inadequately reflect a test suite’s ability to validate program behavior, while mutation testing incurs prohibitively high computational overhead. To address this, we propose Metamorphic Coverage (MC), a novel coverage metric tailored for mutation testing. MC quantifies test sensitivity to potential faults by measuring path divergence between metamorphic input pairs. It is the first systematic formulation and empirical evaluation of a coverage criterion explicitly designed for metamorphic relations. We validate MC across diverse systems—including databases, compilers, and constraint solvers. Experiments on 64 real-world defects show MC achieves a fault-localization coverage of 78.1%, with sensitivity four times that of line coverage and mean overhead only one-sixth thereof. Moreover, MC’s computational cost is merely 0.28% of that of full mutation testing, while improving defect detection rate by 41% over conventional coverage metrics.
To address the lack of reliable and rigorous adequacy criteria for metamorphic testing (MT), this paper establishes, for the first time, an attribute-driven theoretical framework for MT adequacy grounded in the **essential properties** of the system under test. We propose a **computable adequacy metric** that jointly considers metamorphic relation (MR) coverage and the diversity/representativeness of source inputs, enabling objective, quantitative assessment of test quality. Unlike conventional approaches that focus solely on MR coverage or input count, our method holistically integrates semantic correctness with input quality. Empirical evaluation demonstrates that the proposed metric exhibits a strong positive correlation with fault-detection capability: test suites achieving higher adequacy yield an average 23.6% improvement in fault detection rate. Moreover, the metric effectively guides test suite optimization.
This work addresses the exacerbation of the oracle problem in software testing caused by the generative and open-ended nature of large language models (LLMs), which traditional methods struggle to handle. Through a systematic review of 93 studies, it proposes and constructs a novel bidirectional synergy framework between metamorphic testing (MT) and LLMs. On one hand, MT is leveraged to evaluate LLM behaviors concerning hallucination, fairness, and robustness; on the other, LLMs’ semantic understanding and code generation capabilities are harnessed to automate key MT tasks—namely, metamorphic relation discovery, input transformation, and test execution. The study establishes a unified taxonomy encompassing both “MT for LLMs” and “LLMs for MT,” thereby providing a structured foundation and a co-evolutionary pathway for AI quality assurance.
This work addresses the limitation of traditional delta debugging, which relies on test oracles to verify whether reduced inputs preserve the target property—a requirement that hinders its applicability in oracle-absent scenarios. To overcome this challenge, the paper introduces metamorphic testing into the delta debugging pipeline for the first time, constructing an oracle-free validation function that replaces the original oracle-dependent testing mechanism. The proposed approach, termed DDMT, seamlessly integrates with various delta debugging algorithms. Empirical evaluation across 66 subjects demonstrates that DDMT not only maintains or even improves reduction effectiveness and query efficiency but also significantly extends the applicability of delta debugging to environments lacking test oracles.
This work addresses the limitation of conventional execution coverage in UI component testing, which fails to verify whether tests adequately capture behavioral relationships implied by APIs and documentation. The paper proposes the first evaluation framework based on inferred metamorphic relations (MRs): it automatically derives MRs using a UI-specific taxonomy from source code and documentation, aligns test executions to these MRs through deterministic and semantic analysis, and introduces relation-level MR coverage as a novel metric. By treating inferred MRs as empirical benchmarks for behavioral validation, the approach exposes verification gaps invisible to traditional coverage metrics—particularly in weak-oracle scenarios. Empirical results across three LLM configurations show MR coverage ranging only from 42.5% to 47.6%, substantially lower than MR reachability; uncovered MRs are predominantly of the weak-oracle type, demonstrating that MR coverage meaningfully complements conventional metrics and offers practical utility in fault detection and issue mapping.
Existing compiler testing techniques are often ill-suited for transpilers, as they typically lack multiple equivalent implementations and may produce non-executable output code. This work introduces metamorphic testing to transpiler validation by proposing the notion of “mutation consistency”: it defines metamorphic relations at the source-code level to verify whether structurally consistent and expected changes manifest in the generated code when the input DSL program undergoes semantics-preserving mutations. This approach enables defect detection without requiring execution of the generated code. We develop a mutation-based modeling method for metamorphic relations, a source-level structural consistency analysis mechanism, and implement an automated tool, MCP-Tester. Evaluated on real-world technology migration cases, our method effectively uncovers transpiler bugs that elude conventional fuzzing approaches.
This work addresses the challenges in testing augmented reality (AR) applications, where dynamic virtual–physical interactions complicate oracle definition and manual metamorphic relation (MR) construction is costly. To overcome these issues, the authors propose a novel approach that automatically generates and refines MRs by integrating repository-level code context with a multi-agent negotiation mechanism. For the first time, large language model–based reasoning is orchestrated with contextual awareness at the repository scale to enhance MR coverage and reduce redundancy. Evaluation on 142 mobile AR projects shows that hierarchical context modeling covers 7,004 out of 14,916 candidate MRs; 88.2% of the refined MRs are contextually relevant and conflict-free. Manual validation confirms that the generated MRs are logically sound, executable, and successfully detect real nonequivalent bugs.
This work addresses the lack of systematic validation for unit tests generated by foundation models, which hinders reliable assessment of their correctness, utility, and maintainability. To bridge this gap, we introduce TestMap—an open-source infrastructure tailored for C#/.NET projects—that establishes an evidence-centered framework for automated test generation across its entire lifecycle, encompassing mapping, execution, repair, evaluation, and experiment tracking. By integrating repository analysis, source-test mapping, coverage and mutation testing, static analysis, test smell detection, and model-guided generation, TestMap enables observable, reproducible, and comparable experimentation across diverse models, prompts, and strategies. The framework further uncovers limitations of current models, missing contextual information, repair overhead, and latent defects in the system under test, thereby providing an empirical foundation for trustworthy test generation.