Score
Designs, implements, and analyzes verification environments and assertion-based checks for register-transfer-level (RTL) hardware designs using SystemVerilog; this includes writing SystemVerilog Assertions (SVA), integrating SVA into simulation and formal verification flows, and using assertions, testbenches, coverage, and debug techniques to validate RTL correctness against specifications.
In RTL design, root-cause analysis and repair of SystemVerilog Assertion (SVA) failures heavily rely on expert knowledge, with minimal automation support. This paper introduces AssertSolver—the first open-source domain-specific large language model (DS-LLM) tailored for SVA debugging. Our method comprises two key innovations: (1) a domain-knowledge-enhanced architecture integrating RTL semantics, assertion logic, and temporal relationships; and (2) an error-driven synthetic data generation and reinforcement fine-tuning paradigm, significantly improving attribution accuracy and fix generation for assertion violations. Evaluated on a novel, comprehensive benchmark we curate, AssertSolver achieves 88.54% pass@1 bug-fixing accuracy—outperforming OpenAI o1-preview by 11.97%. To foster reproducibility and community advancement, we fully open-source the model, training data, evaluation benchmark, and implementation code.
Existing LLM-based assertion generation methods struggle to model cross-layer semantic correlations between design specifications and RTL code, resulting in low assertion coverage and poor accuracy. To address this, we propose a cross-layer signal bridging mechanism that integrates chain-of-thought reasoning with signal mapping analysis to explicitly align natural-language specifications with RTL signal semantics, enabling precise, automated SystemVerilog assertion generation. Our end-to-end framework synergistically combines large language models, formal verification, and mutation testing feedback to significantly enhance assertion completeness and verifiability. Experimental evaluation demonstrates that our approach outperforms state-of-the-art methods across key metrics—including formal verification pass rate, cone-of-influence coverage, proof-core coverage, and mutation kill rate—establishing new performance benchmarks for specification-driven assertion synthesis.
This work addresses the limitations of existing approaches to automatic SystemVerilog Assertion (SVA) generation—namely, erroneous signal references, missing temporal constraints, and the absence of formal correctness guarantees—by introducing ProofLoop, a tool-augmented ReAct agent that integrates a formal verification solver into the large language model reasoning loop. ProofLoop employs a two-stage pipeline: it first retrieves design context using EDA tools and an AST-based vector database, then iteratively refines assertions through structural queries in JasperGold and multi-round feedback from formal verification. Evaluated on the FVEval Design2SVA benchmark, ProofLoop achieves 93.7% syntactic correctness and 82.0% functional correctness. Ablation studies confirm that each component contributes significantly and orthogonally to overall performance.
This work addresses the labor-intensive and error-prone process of manually crafting SystemVerilog Assertions (SVA) from specification documents in assertion-based verification (ABV). It presents the first systematic analysis of the key challenges involved in leveraging large language models (LLMs) for automated SVA generation and proposes a principled methodology that ensures high-quality, standardized outputs. By integrating natural language processing with formal verification techniques, the study formulates guiding principles and practical strategies tailored to the generation of reliable and precise assertions. This approach provides both theoretical grounding and a viable pathway toward building efficient, robust automated verification workflows.
SystemVerilog assertions in RTL verification are error-prone and difficult to maintain, while existing automated assertion generation methods lack repair capabilities. Method: This paper proposes AssertFix—the first end-to-end, LLM-based assertion auto-repair framework—integrating RTL semantic understanding, precise error localization, root-cause diagnosis, and fine-grained error classification, enabling customizable repair strategies and minimizing manual intervention. Unlike conventional approaches that treat assertion generation as a post-verification task, AssertFix achieves full automation from error identification to correction. Contribution/Results: Evaluated on the OpenCores benchmark, AssertFix improves assertion repair rate by 42.6% and increases verification coverage by 18.3% on average, demonstrating its effectiveness, robustness, and practical engineering applicability.
This work addresses the heavy reliance on manual modeling and proof effort in formal verification of SystemVerilog RTL designs by proposing the first fully automated framework for translating RTL to Lean 4. The approach introduces a four-layer hierarchical theorem library encompassing combinational logic, sequential updates, single-cycle behaviors, and reachability/invariant properties. It further integrates an LLM-driven proof loop that automatically generates intermediate lemmas, admitting only those formally verified by the Lean kernel into a reusable lemma pool. Evaluated on six designs, the method successfully produced 403 theorems, of which 287 foundational lemmas were automatically reusable, achieving a reuse rate of 80.2%. This significantly enhances the automation and scalability of formal RTL verification.
Large language models (LLMs) are increasingly being explored for automating SystemVerilog Assertion (SVA) generation, yet most evaluations report correctness on a single syntactic representation of an input. Such point accuracy does not reveal whether a model's correct output is stable when the same RTL behavior is written differently. This paper presents a controlled metamorphic evaluation of LLM-based SVA generation under semantics-preserving RTL transformations. Starting from the VERT dataset, we construct a quality-filtered conditional-control pool and a stratified 40-program evaluation set containing 295 assignment behaviors. We evaluate two open code models, Qwen2.5-Coder-7B and DeepSeek-Coder-V2-Lite, with an identical evaluation prompt and greedy decoding. Three transformations are studied: operand reordering, deterministic identifier renaming, and redundant parenthesization. Beyond baseline and transformed accuracy, we measure conditional robustness, invariance failure, and any-flip rate, with 10,000-sample clustered bootstrap intervals at the RTL-program level. Across all six model-transformation conditions, 9.7%-27.0% of behaviors that were correct on the original RTL become incorrect after a semantics-preserving transformation. Aggregate accuracy can therefore hide substantial instability: under identifier renaming, DeepSeek-Coder-V2-Lite improves from 53.9% to 63.7% accuracy while 19.5% of its originally correct behaviors fail. Manual review of 30 sampled correct-to-wrong transitions identifies dropped path predicates, branch-polarity errors, Boolean-structure corruption, and output-contract violations. The results show that point accuracy alone is insufficient for characterizing LLM reliability in assertion generation and motivate robustness-aware evaluation for AI-assisted hardware verification.
为解决硬件功能验证中难以生成全面可靠的断言集问题,提出一种结合形式化探索与神经符号细化的覆盖驱动RTL断言生成框架。
This study addresses the challenges of semantic alignment and tool integration in automating hardware verification, particularly concerning assertion generation, debugging, and formal reasoning. To overcome these limitations, this work proposes a neuro-symbolic hybrid architecture that leverages large language models (LLMs) as core orchestration components. By integrating prompt engineering, retrieval-augmented generation, agentic workflows, and SAT/SMT solver optimization techniques, the proposed framework establishes semantic consistency as a critical breakthrough for automated verification pipelines. Furthermore, this paper systematically reviews the application paradigms of LLMs across the entire hardware verification workflow and empirically validates the effectiveness of the hybrid architecture. Finally, it provides an in-depth analysis of the limitations inherent in current evaluation methodologies and outlines promising directions for future research in AI-driven electronic design automation.
本文提出了一种基于大型语言模型的行为驱动硬件开发流程,通过定义形式验证Gherkin场景来减少自然语言规范的模糊性,提高硬件设计的形式验证效果。