Score
Designs and analyzes plans, strategies, and activities to reach target coverage levels in a verification or testing effort, including creation of coverage closure plans and prioritization of coverage gaps. Builds and evaluates tests, stimulus generation (drive coverage), and post-run coverage analyses to close functional and other coverage holes and quantify remaining verification risk.
Traditional formal verification methods often struggle to achieve full coverage within project timelines due to difficulties in coverage convergence. This work proposes the first formal verification workflow that integrates autonomous AI agents with generative AI, leveraging large language models to automatically analyze coverage gaps and generate corresponding formal properties, thereby establishing a closed-loop automation for coverage analysis and property generation. Evaluated on both open-source and internal designs, the approach significantly improves coverage metrics, with greater gains observed as design complexity increases, demonstrating its effectiveness and scalability.
This work addresses the limitations of traditional structural coverage metrics in embedded software testing, which are often confined to the unit level and fail to reflect true coverage completeness in integration and system testing. Instrumentation-based approaches risk perturbing runtime behavior, while pure tracing techniques suffer from unreliability under high compiler optimization. To overcome these challenges, the paper proposes an integration-test-driven coverage strategy featuring a novel “integration-first” closed-loop workflow. By synergistically combining embedded tracing with hybrid runtime analysis (hRA) to preserve semantic boundaries, and leveraging source-to-target mapping for evidential traceability alongside Hyper Coverage for cross-variant merging, the approach establishes a unified evidence-integration mechanism. Evaluated on -O3-optimized release binaries, it reliably achieves branch, condition, and MC/DC coverage measurements and precisely identifies source code lines consistently uncovered across all variants, thereby significantly enhancing confidence in the test completeness of embedded systems.
This work addresses a critical limitation in existing cloud skill testing, which focuses solely on task success rates and fails to reveal uncovered behaviors, leaving test adequacy unquantifiable. To bridge this gap, the study introduces the first formal definition of test coverage units, coverage relationships, and a computational pipeline for cloud skills. It proposes a coverage evaluation method grounded in natural language skill packages: by parsing user prompts and initial resource states, the approach reconstructs operational obligations and models workflow context to establish an end-to-end measurement pipeline. Integrating model-assisted candidate generation with expert review, the framework produces auditable coverage reports and closes the loop by mapping coverage gaps to source-level test improvement recommendations. Empirical results demonstrate that this methodology enables quantitative assessment of test coverage for production-grade cloud skills, substantially enhancing their reliability and observability.
Ensuring trustworthiness across the full lifecycle—design, operation, and evolution—of autonomous systems faces two key challenges: the fragmentation between design-time and run-time assurance, and insufficient adaptability to dynamic environmental and operational changes. This paper proposes a unified continuous assurance framework that integrates formal verification (via RoboChart), probabilistic risk analysis (using PRISM), and assurance case modeling through a model-driven approach, enabling co-modeling and dynamic updating of assurance artifacts. The framework supports automated traceability, reconstruction, and regeneration of assurance arguments, thereby establishing an end-to-end trust chain spanning design, operation, and evolution phases. An Eclipse plugin implements automated model transformation and argument generation. Evaluated on a nuclear inspection robot case study, the framework significantly enhances system trustworthiness, regulatory compliance, and alignment with tripartite AI principles—accountability, transparency, and robustness.
In hardware logic verification, conventional dynamic simulation and module-level approaches fail to ensure comprehensive transaction-level functional coverage and suffer from poor verification result reusability. This paper proposes the first transaction-level (TL) hierarchical deductive formal verification framework. It extends the PDVL language to support functional coverage and assertion modeling, compiles PDVL specifications into Coq-verifiable Gallina code, and automates the translation of functional coverage goals into proof obligations. The framework enables cross-layer verification reuse and supports formal verification of SVA assertions. Crucially, achieving 100% functional coverage is formally equivalent to establishing system completeness—a rigorous proof of correctness. Our approach delivers high verification accuracy while significantly improving reusability and efficiency over traditional assertion-based verification (ABV) and simulation hybrid methods.
This work addresses the inefficiency of coverage convergence in hardware verification, where existing large language model (LLM) agents lack systematic analysis of hard-to-reach coverage holes and principled mechanisms for allocating reasoning resources. The authors propose a two-tier agent framework that integrates a base Codex agent with a domain-enhanced LangGraph system, establishing the first taxonomy of coverage gaps—categorized by methodological ceilings and reasoning frontiers—to expose fundamental limitations of purely LLM-driven verification. They introduce a profile-driven agent design paradigm, incorporating fine-grained tracking of token consumption across six categories and a coverage feedback loop. Evaluated on multiple designs, the approach achieves 95–99% coverage while reducing token usage by 4–13× and accelerating convergence by 2–4× compared to general-purpose baselines.
This work addresses the limitations of traditional RTL functional verification, which heavily relies on manual effort, and overcomes the shortcomings of existing LLM-based approaches that suffer from context fragmentation, leading to interface mismatches and coverage metrics decoupled from specifications. The paper proposes a novel agent-based framework that integrates an execution control layer, an evolvable knowledge system, and specification-anchored coverage modeling to establish an end-to-end automated verification loop. This framework enables, for the first time, fully automatic generation of complete verification environments without human intervention and precisely links each coverage bin to specific specification behaviors, facilitating diagnosable gap identification and repair. Evaluated on eight RTL designs, the approach achieves 100% success in verification environment generation, with average line, branch, toggle, and functional coverage rates of 98.4%, 97.2%, 97.0%, and 83.2%, respectively.
This work addresses the inadequacy of traditional code coverage criteria in guiding prompt-centric testing within large language model (LLM)-driven software development. It proposes a novel prompt-level coverage metric grounded in LLM attention mechanisms, shifting the focus of coverage adequacy from source code to natural language prompts. By quantifying how well test cases satisfy the requirements expressed in prompts, the method directs the generation of more effective tests. Empirical evaluation across multiple LLMs and datasets demonstrates that this approach significantly outperforms conventional code coverage techniques, uncovering over 30% more defects on average. The study thus establishes a foundational testing metric tailored to the emerging paradigm of LLM-based programming.