Score
Design and define quantitative metrics and associated instrumentation that measure how thoroughly tests or verification suites exercise specified behaviors, code paths, and assertions (including functional coverage points, assertion-based coverage, and code/test coverage metrics). Build functional coverage models and tooling, collect and analyze coverage data and reports, and plan or optimize tests and instrumentation to close coverage gaps and satisfy the defined metrics.
This work addresses the limitations of traditional structural coverage metrics in embedded software testing, which are often confined to the unit level and fail to reflect true coverage completeness in integration and system testing. Instrumentation-based approaches risk perturbing runtime behavior, while pure tracing techniques suffer from unreliability under high compiler optimization. To overcome these challenges, the paper proposes an integration-test-driven coverage strategy featuring a novel “integration-first” closed-loop workflow. By synergistically combining embedded tracing with hybrid runtime analysis (hRA) to preserve semantic boundaries, and leveraging source-to-target mapping for evidential traceability alongside Hyper Coverage for cross-variant merging, the approach establishes a unified evidence-integration mechanism. Evaluated on -O3-optimized release binaries, it reliably achieves branch, condition, and MC/DC coverage measurements and precisely identifies source code lines consistently uncovered across all variants, thereby significantly enhancing confidence in the test completeness of embedded systems.
Traditional test adequacy metrics, such as code coverage and mutation testing, focus on implementation details and struggle to capture discrepancies between expected and actual program behavior. This work proposes an automated approach that extracts method-level expected behaviors from natural language documentation and source code, then maps them to existing test cases, thereby formalizing and empirically evaluating “behavioral gaps”—a dimension of test adequacy independent of structural metrics. By integrating natural language processing, static analysis, and behavioral mapping techniques, our method identifies 20,729 behaviors across ten Java open-source libraries with 93.1% precision, revealing that 17.5% of expected behaviors remain entirely untested. Notably, these gaps persist even in methods exhibiting high code coverage or high mutation kill rates, exposing a systematic deficiency in current testing practices—including automatically generated tests—in validating intended program behavior.
This study challenges the common assumption that high code coverage implies high fault localization accuracy in spectrum-based fault localization (SBFL). We systematically evaluate the effectiveness of automatically generated test cases—produced by tools such as EvoSuite and Randoop—against manually written tests, measuring SBFL score (our primary quality metric), mutation kill rate, and branch coverage across 42 real-world defects from Defects4J. Contrary to expectations, although automatically generated tests achieve 18% higher average branch coverage, they yield 23% lower SBFL scores. Moreover, in deeply nested code regions, their fault localization precision degrades by up to 41%. These findings reveal a fundamental inconsistency between coverage metrics and actual fault localization capability. The work provides the first empirical evidence centered on SBFL score as the key evaluation criterion, offering critical support for hybrid testing strategies and advocating a paradigm shift in test quality assessment—from coverage-oriented to localization-oriented evaluation.
This work proposes NQC2, a non-intrusive code coverage collection mechanism based on QEMU plugins, designed to address the challenge of applying traditional coverage analysis—typically reliant on operating systems and file systems—to bare-metal embedded programs. By leveraging dynamic binary translation, NQC2 extracts execution path information from within QEMU during emulation and saves it directly to the host machine, without requiring modifications to the target program or a customized QEMU build. This approach enables, for the first time, zero-instrumentation coverage analysis for bare-metal embedded systems. Experimental results demonstrate that NQC2 achieves up to an 8.5× performance improvement over Xilinx’s comparable solution, significantly enhancing both the efficiency and applicability of testing for embedded software.
This work addresses the challenge of runtime code coverage analysis in AAA-grade C++ game engines, where traditional whole-program instrumentation incurs prohibitive performance overhead and compromises test stability. To overcome these limitations, the authors propose a developer-submission-oriented selective instrumentation approach that leverages a lightweight, incremental coverage collection mechanism. Integrated with a customized compiler toolchain and an industrial-scale testing pipeline, this method substantially reduces both compilation and runtime overhead. Evaluated across more than 2,000 code submissions, it maintains minimal compilation cost, sustains frame rates above 50% of baseline even in worst-case scenarios, and introduces no failures in automated tests—demonstrating, for the first time, a practical and scalable solution for real-time coverage analysis in large-scale game engines.
This work addresses a critical limitation in existing cloud skill testing, which focuses solely on task success rates and fails to reveal uncovered behaviors, leaving test adequacy unquantifiable. To bridge this gap, the study introduces the first formal definition of test coverage units, coverage relationships, and a computational pipeline for cloud skills. It proposes a coverage evaluation method grounded in natural language skill packages: by parsing user prompts and initial resource states, the approach reconstructs operational obligations and models workflow context to establish an end-to-end measurement pipeline. Integrating model-assisted candidate generation with expert review, the framework produces auditable coverage reports and closes the loop by mapping coverage gaps to source-level test improvement recommendations. Empirical results demonstrate that this methodology enables quantitative assessment of test coverage for production-grade cloud skills, substantially enhancing their reliability and observability.
This work addresses the challenge of balancing software quality, testability, and maintainability under rapid iteration and frequent requirement changes. It proposes Algorithm-Driven Development (ADD), a novel approach that unifies requirements specification and technical design by using algorithm flowcharts as a single, coherent artifact. This integration enables end-to-end modeling of requirements, architecture, and testing. Leveraging this model, the system automatically generates high-coverage acceptance tests and incorporates continuous integration with code coverage feedback. Industrial adoption at Dassault Systèmes demonstrates that ADD achieves over 95% code coverage, substantially reduces defect density, and ensures a stable delivery cadence, outperforming conventional test-driven development and test-after approaches.
This work addresses the inadequacy of traditional code coverage criteria in guiding prompt-centric testing within large language model (LLM)-driven software development. It proposes a novel prompt-level coverage metric grounded in LLM attention mechanisms, shifting the focus of coverage adequacy from source code to natural language prompts. By quantifying how well test cases satisfy the requirements expressed in prompts, the method directs the generation of more effective tests. Empirical evaluation across multiple LLMs and datasets demonstrates that this approach significantly outperforms conventional code coverage techniques, uncovering over 30% more defects on average. The study thus establishes a foundational testing metric tailored to the emerging paradigm of LLM-based programming.
Existing testing approaches struggle to effectively detect functional bugs that do not cause program crashes: manually written unit tests are costly, heuristic-based test generation lacks semantic understanding, and fuzzing relies heavily on crash signals. To address this limitation, this work proposes LISA, a novel framework that uniquely integrates large language models (LLMs) with program invariants. LISA employs semantically guided API call sequence generation and an API n-gram–based feedback mechanism to iteratively refine test cases. The approach substantially improves both the detection rate and precision of functional defects, outperforming state-of-the-art fuzzing techniques and LLM-driven testing methods in terms of code coverage and the generation of high-confidence bug reports.
Existing research often relies on code coverage and mutation score as proxy metrics to evaluate the effectiveness of test cases generated by large language models (LLMs), yet their correlation with actual fault-detection capability remains unclear. This study conducts a large-scale empirical analysis of test suites produced by diverse LLMs across varied testing scenarios, systematically examining the relationships among coverage, mutation score, and real defect detection performance. The findings reveal that the validity of these proxy metrics is highly context-dependent: they offer some predictive value in regression testing but prove unreliable when the target code contains faults. Furthermore, the size of the test suite has limited influence on these correlations. These results challenge conventional assumptions in test evaluation and provide new empirical grounding for assessing LLM-generated tests.