Score
Design verification is the competence to design, build, and execute verification artifacts and processes that demonstrate a design satisfies its specifications; this includes creating verification plans, testbenches, simulation and formal models, assertion suites, test harnesses, and coverage metrics. It also encompasses analyzing verification results to detect, localize, and quantify design defects and to iterate on the design or verification strategy.
This work addresses the labor-intensive and error-prone process of manually crafting SystemVerilog Assertions (SVA) from specification documents in assertion-based verification (ABV). It presents the first systematic analysis of the key challenges involved in leveraging large language models (LLMs) for automated SVA generation and proposes a principled methodology that ensures high-quality, standardized outputs. By integrating natural language processing with formal verification techniques, the study formulates guiding principles and practical strategies tailored to the generation of reliable and precise assertions. This approach provides both theoretical grounding and a viable pathway toward building efficient, robust automated verification workflows.
Existing hardware assertion generation methods suffer from poor scalability to industrial-scale designs, low assertion quality, insufficient functional coverage depth, and limited interpretability. Method: We introduce the first LLM evaluation benchmark for Verilog designs—comprising 100 open-source circuits and formally verified “gold-standard” assertions—and propose a dedicated evaluation framework integrating functional equivalence checking, multi-dimensional quality metrics, and context-example sensitivity analysis. Contribution/Results: This work fills a critical gap in quantitative LLM assessment for hardware verification. Experiments reveal that state-of-the-art LLMs achieve less than 35% assertion correctness overall, though performance improves markedly with increasing context examples. Systematic deficiencies are identified in modeling temporal logic and finite-state machines. Our benchmark, methodology, and empirical findings provide foundational resources and evidence for advancing LLM-driven hardware verification.
This work addresses the error-prone and labor-intensive process of manually translating natural language specifications into formal representations for chip design verification. It presents the first end-to-end agent framework capable of automatically converting industrial-grade DRAM standard specifications from natural language into the domain-specific language DRAMPyML, which is then seamlessly integrated into hardware verification workflows to generate SystemVerilog assertions, stimuli, and functional coverage metrics. The approach is validated on real-world DRAM specifications, supported by a newly constructed evaluation benchmark, DRAMBench, and an open-sourced dataset to advance research in automated formalization of hardware specifications.
Despite its efficacy in isolated projects, deductive verification has yet to achieve broad industrial adoption. To identify root barriers and key enablers, this paper conducts semi-structured interviews with 30 practitioners, followed by thematic analysis. We systematically uncover fundamental obstacles—including high proof maintenance overhead, limited automation, poor tool usability, and lack of workflow integration—as well as critical enabling factors. Diverging from prior work, we empirically establish *usability* and *workflow adaptability* as core dimensions governing adoption. Based on these findings, we propose three actionable improvement principles: (1) enhancing automation support for proof construction and evolution; (2) reducing proof maintenance burden through modularization and abstraction; and (3) deepening integration with IDEs and CI/CD pipelines. Our empirically grounded insights provide concrete, evidence-based guidance for tool developers, practitioners, and researchers—bridging the gap between academic verification techniques and engineering practice.
Existing code-level formal verification tools scale poorly to large-scale software, while mainstream unit-level verification relies heavily on manual effort, often missing critical defects. This paper proposes the “Unit Proof Framework” research agenda—the first systematic definition of a unit verification paradigm supporting automated decoupling and independent verification of code units. Methodologically, it integrates formal verification, program analysis, modular verification, and automated toolchain design, with deep alignment to industrial development practices (e.g., AWS workflows). Its core contributions include: (1) establishing a scalable, engineering-friendly unit verification methodology; (2) characterizing a taxonomy of key technical challenges; (3) overcoming bottlenecks inherent in manual verification; and (4) significantly improving early detection of code-level defects. Collectively, this work lays the theoretical foundation and provides a practical technical pathway for building high-assurance, deployable automated verification infrastructure.
This work addresses the inefficiency and error-proneness of manually writing SystemVerilog Assertions (SVA) and the lack of realistic verification scenarios in existing large language model (LLM) evaluation benchmarks. To bridge this gap, the paper introduces AssertLLM2—the first open-source benchmark specifically designed for hardware verification assertion generation, encompassing 83 real-world design cases and supporting two critical tasks: bug-prevention and bug-hunting. AssertLLM2 innovatively employs systematically mutated erroneous RTL as input and establishes a multidimensional evaluation framework that assesses syntactic correctness, formal provability, coverage, and mutation detection capability. Built upon authentic design specifications, structured requirement descriptions, golden reference implementations, and fault injection techniques, AssertLLM2 provides a rigorous and comprehensive baseline for evaluating LLM performance in hardware verification.
Traditional simulation struggles to cover rare corner-case scenarios, while formal verification is hindered by limited scalability and high usability barriers. To address these challenges, this work proposes Forbench—a word-level symbolic simulation framework that seamlessly integrates symbolic execution into conventional RTL simulation workflows. By leveraging an SMT solver to support symbolic signals and state transitions, Forbench enables systematic exploration of design behaviors while preserving the semantics of traditional simulation. It further provides a simulation-like Python interface for defining constraints, enabling co-simulation, and performing property checking. Forbench significantly improves verification efficiency without compromising coverage and substantially lowers the barrier to adopting formal methods.
This work addresses the challenge of efficiently verifying large-scale RTL designs generated by high-level synthesis (HLS), which often overwhelm conventional model checking techniques. The authors propose a novel method that leverages high-level semantic information from HLS to automatically generate guided invariants, which augment assertions to accelerate formal verification. A proof-guided selection mechanism is introduced to iteratively refine and identify an optimal set of assertions. This approach represents the first systematic integration of HLS-level features into automated invariant generation, substantially improving verification efficiency. Experimental results across multiple HLS benchmarks demonstrate an average speedup of 2.23×, with a maximum acceleration of 6.05× compared to baseline methods.
This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.