Score
Designs, implements, and executes test suites, frameworks, and procedures to verify that software or systems interoperate correctly across specified operating systems, hardware devices, firmware, and runtime environments; this includes creating test matrices, automated and manual test cases, and harnesses to reproduce cross‑platform behaviors. Analyzes compatibility failures, diagnoses environmental causes (API/ABI differences, drivers, configuration), and produces verification evidence and remediation recommendations.
This study addresses the imbalance in the test pyramid—characterized by an overreliance on coarse-grained integration and system tests, which leads to difficulties in fault localization and slow execution—by proposing, for the first time, a method to automatically generate unit tests from existing integration tests. The approach combines static and dynamic analysis to automatically isolate component dependencies and enhance coverage at the unit level. Implemented as a Node.js tool and evaluated on twelve open-source JavaScript projects, the technique produces high-quality unit tests that significantly improve test suite structure, thereby increasing both testing efficiency and maintainability.
This work addresses the inefficiencies and semantic inconsistencies arising from separately implementing driver and monitor programs in traditional hardware module testing. To overcome this, the authors propose a domain-specific language (DSL) tailored to hardware communication protocols, which enables the unified specification of both driver and monitor logic through an imperative syntax, thereby ensuring their semantic consistency for the first time. Building upon this DSL, they develop a prototype tool that leverages waveform parsing and transaction-level trace inference techniques to accurately reconstruct protocol-compliant transaction sequences from raw signal waveforms. Experimental results demonstrate that the approach significantly improves development efficiency, with further validation planned on real-world interconnect protocols such as Wishbone and AXI-Stream.
To address the challenges of automating full-system integration testing, the limited coverage of hardware-software co-faults by conventional formal verification, and the high cost of system performance characterization, this paper proposes TestIt—the first lightweight Software-Based Self-Testing (SBST) integration testing framework tailored for RTL-level systems. TestIt supports dual execution environments: cycle-accurate simulation and FPGA deployment (on PYNQ-Z2), dynamically generating test vectors and golden reference outputs to enable efficient CI/CD integration in open-source RTL development. Innovatively extending SBST to RTL system verification, TestIt detects hardware-software interaction bugs often missed by formal methods and enables low-overhead, rapid system performance profiling. Evaluated on the X-HEEP RISC-V microcontroller, TestIt achieves an 11× speedup on FPGA over RTL simulation and successfully uncovers multiple co-faults undetected by formal verification.
This work proposes NQC2, a non-intrusive code coverage collection mechanism based on QEMU plugins, designed to address the challenge of applying traditional coverage analysis—typically reliant on operating systems and file systems—to bare-metal embedded programs. By leveraging dynamic binary translation, NQC2 extracts execution path information from within QEMU during emulation and saves it directly to the host machine, without requiring modifications to the target program or a customized QEMU build. This approach enables, for the first time, zero-instrumentation coverage analysis for bare-metal embedded systems. Experimental results demonstrate that NQC2 achieves up to an 8.5× performance improvement over Xilinx’s comparable solution, significantly enhancing both the efficiency and applicability of testing for embedded software.
Binary program symbolic execution suffers from semantic distortion and implementation errors introduced during intermediate representation (IR) translation. Method: This paper proposes the first instruction-level symbolic execution framework directly grounded in formal ISA semantics (Rock/Sail), bypassing conventional IR abstractions by compiling machine-readable ISA specifications into SMT-solvable symbolic semantic models and integrating them into a binary analysis platform. Contributions/Results: (1) The first end-to-end automated pipeline from formal ISA semantics to symbolic execution; (2) Demonstrated scalability on RISC-V—modeling new instructions requires only a few hours; (3) Discovered five previously unknown ISA semantic implementation bugs in angr; (4) Achieved high-fidelity branch modeling and solving capability. The framework significantly improves the accuracy, trustworthiness, and development efficiency of binary symbolic execution.
This work addresses the challenge of applying equivalence class partitioning—a testing requirement under ISO 26262—to legacy embedded firmware in the absence of complete specification documents. The authors propose a binary-level method that automatically infers output-oriented equivalence classes by reconstructing control flow and performing guided symbolic execution to analyze function behavior. Execution paths are clustered based on observable outputs, such as return values and output parameters, and the resulting equivalence classes are represented in a human-readable form to support test design. To the best of the authors’ knowledge, this is the first approach capable of inferring equivalence classes directly from binaries without source code or documentation for safety-critical embedded software. Industrial case studies demonstrate that the inferred classes align closely with expert expectations and offer both high readability and practical utility, effectively aiding functional comprehension and compliance testing of legacy firmware.
This work addresses the limitations of traditional structural coverage metrics in embedded software testing, which are often confined to the unit level and fail to reflect true coverage completeness in integration and system testing. Instrumentation-based approaches risk perturbing runtime behavior, while pure tracing techniques suffer from unreliability under high compiler optimization. To overcome these challenges, the paper proposes an integration-test-driven coverage strategy featuring a novel “integration-first” closed-loop workflow. By synergistically combining embedded tracing with hybrid runtime analysis (hRA) to preserve semantic boundaries, and leveraging source-to-target mapping for evidential traceability alongside Hyper Coverage for cross-variant merging, the approach establishes a unified evidence-integration mechanism. Evaluated on -O3-optimized release binaries, it reliably achieves branch, condition, and MC/DC coverage measurements and precisely identifies source code lines consistently uncovered across all variants, thereby significantly enhancing confidence in the test completeness of embedded systems.
This study addresses the unclear relationships among information sources, generation strategies, and quality evidence in test case generation using large language models (LLMs). Through a systematic literature review of 95 studies, this work constructs a multidimensional taxonomy and a benchmark analysis framework. Specifically, it proposes a four-dimensional classification system that elucidates how execution feedback influences oracle independence. Furthermore, it establishes a unified theoretical framework connecting the generation process with quality assessment. By identifying independent oracle evaluation as a critical yet underexplored dimension, this research formulates a future agenda centered on rigorous, oracle-independent quality measurement for LLM-generated test cases.
该研究提出一种基于大语言模型和数据手册的框架,用于早期嵌入式系统设计中的硬件兼容性验证,通过构建设计图和分解任务提高准确性和效率。
This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.