Score
Design, implement, and optimize register‑transfer‑level (RTL) hardware descriptions by writing synthesizable SystemVerilog RTL code, performing simulation and verification, and applying RTL debug methodologies to find and fix functional and timing issues. Drive synthesis, place‑and‑route and RTL‑to‑GDS flows and iteratively refine gate‑level rewrites to meet power/performance/area (PPA) targets, using PPA tools and quantitative metrics to guide changes and optionally incorporating automated or LLM/agentic-in-the-loop assistance for optimization and design iteration.
To address the bottleneck in VLSI design where RTL-stage PPA (power, performance, area) estimation relies on time-consuming full synthesis—hindering rapid iteration—this paper proposes the first pre-synthesis machine learning framework operating directly on HDL source code. Our method introduces a bit-level Simple Operation Graph (SOG) representation that explicitly models the semantic mapping between RTL constructs and post-synthesis structures. We further design a standard-cell-library-aware tree-based architecture enabling end-to-end PPA prediction solely from Verilog code and library files, without requiring toggle information or synthesis intermediates. Evaluated on 147 industrial-scale RTL designs, our framework achieves 98% accuracy for worst negative slack (WNS), 98% for total negative slack (TNS), and 90% for power estimation—significantly outperforming prior approaches. The method demonstrates strong generalization across diverse designs and high engineering practicality.
This work proposes RTL-OPT, a new benchmark addressing the limitations of existing evaluations that primarily focus on syntactic correctness of RTL code while inadequately assessing power, performance, and area (PPA) optimization quality. RTL-OPT comprises 36 handcrafted digital circuit tasks, each paired with unoptimized and expert-optimized RTL implementations. The benchmark introduces an end-to-end automated evaluation pipeline that integrates formal functional equivalence checking with quantitative PPA analysis. It enables the first systematic assessment of large language models’ capabilities in RTL optimization, incorporating optimization patterns reflective of industrial practice and capturing dimensions often overlooked by conventional synthesis tools. RTL-OPT thus provides a standardized, quantifiable platform for evaluating LLM-driven hardware design optimization.
To address key challenges in RTL code optimization—including error-prone manual rewriting, limited capability of traditional compilers in handling complex design constraints, and poor alignment between LLM-generated outputs and user intent—this paper proposes the first neuro-symbolic framework. Our method integrates large language model (LLM)-driven RTL rewriting, abstract syntax tree (AST)-based template retrieval-augmented generation (RAG), and fine-grained finite-state machine (FSM) symbolic analysis supporting state merging and partial reduction. It further incorporates formal equivalence checking and test-driven co-verification. This approach transcends the limitations of pattern-matching compilers: on the RTL-Rewriter benchmark, it achieves 43.9% lower power consumption, 62.5% higher performance, and 51.1% smaller area compared to state-of-the-art methods, as validated by Synopsys Design Compiler and Yosys.
Current large language models (LLMs) lack standardized evaluation for hardware description language (HDL) code generation, particularly for synthesizable, functionally correct communication protocol implementations. Method: We introduce the first protocol-level RTL generation benchmark targeting SPI, I²C, UART, and AXI protocols, featuring multi-abstraction-level generation tasks and a rigorous synthesis-readiness validation pipeline—including syntax checking, logic synthesis, and UVM-driven waveform simulation. Contribution/Results: Evaluating 12 prominent open- and closed-weight LLMs, we find only two models pass all functional correctness checks, with an average synthesis success rate below 35%. Results reveal pervasive deficiencies in protocol-specific timing modeling and concurrent control handling. This benchmark fills a critical gap in evaluating LLMs for digital circuit protocol implementation and establishes a new evaluation paradigm for HDL code generation capability.
This work addresses the effectiveness of large language models (LLMs) in optimizing complex, timing-critical Register Transfer Level (RTL) code expressed in temporal logic. Method: We introduce the first timing-sensitive, four-subset RTL benchmark suite and propose a metamorphic testing framework—based on semantic-preserving code transformations—to systematically quantify LLM capabilities in modeling timing-aware control flow and optimizing across clock domains. Contribution/Results: Experiments show that while LLMs outperform conventional compilers on basic combinational logic optimization, they exhibit significant limitations in handling multi-cycle paths, asynchronous handshaking, and timing constraint propagation—revealing a fundamental deficiency in hardware timing semantics comprehension. This study pioneers the application of formal metamorphic testing to RTL-level LLM evaluation, establishing a reproducible benchmarking methodology and identifying concrete directions for advancing trustworthy AI-assisted hardware design.
This work addresses the challenge that large language models (LLMs) often generate erroneous RTL code due to ambiguous or misinterpreted specifications, with such errors typically surfacing only during simulation and proving difficult to trace. To mitigate this, the paper proposes VeriRefine, a novel approach that first refines informal specifications into explicit Abstract Signal Transition Functions (ASTFs)—serving as a verifiable prelude to RTL generation. The method incorporates a five-layer auditing mechanism to validate design intent across dimensions including completeness, consistency, and FSM integrity, and enables targeted debugging by tracing simulation failures back to either specification misunderstandings or coding errors. Evaluated on RTLLM v2.0 and VerilogEval-Human v2, VeriRefine achieves functional correctness rates of 94.0% and 98.1%, respectively, substantially enhancing the reliability and synthesizability of LLM-generated RTL.
This work addresses the challenge that large language models (LLMs) often introduce semantic or logical errors when generating hardware RTL code, failing to meet the stringent reliability requirements of chip design. To overcome this limitation, the paper proposes a novel hardware generation framework that integrates LLMs with formal methods, uniquely combining LLM-driven iterative refinement with formal verification. The approach leverages predefined transformation rules to guide the LLM in progressively refining high-level specifications into RTL code that is formally verifiable for correctness. This integration enhances both the interpretability and reliability of the code generation process. Experimental results demonstrate that the method is not only effective but also efficient in producing correct RTL implementations, thereby offering a promising pathway toward trustworthy LLM-assisted hardware design.
This work addresses the inefficiency of traditional RTL design flows in optimizing power, performance, and area (PPA), which struggle to automatically explore high-quality implementations. The authors propose an iterative optimization framework that integrates large language model (LLM) agents with circuit-level synthesis. By leveraging a multi-round elite pool mechanism, the approach combines LLM-generated RTL code, gate-level rewriting, and arithmetic architecture search, guided by PPA metrics from Yosys/OpenROAD to iteratively evolve designs. Experimental results on the ASAP7 technology node demonstrate that the method achieves a 35% reduction in area and a 45% decrease in delay for an IEEE-754-compliant 16-bit floating-point multiplier, significantly outperforming reference designs produced by commercial tools across the Pareto front. These findings validate the effectiveness of synergistically coupling high-level intelligent guidance with low-level physical optimization.
This work addresses the heavy reliance on manual modeling and proof effort in formal verification of SystemVerilog RTL designs by proposing the first fully automated framework for translating RTL to Lean 4. The approach introduces a four-layer hierarchical theorem library encompassing combinational logic, sequential updates, single-cycle behaviors, and reachability/invariant properties. It further integrates an LLM-driven proof loop that automatically generates intermediate lemmas, admitting only those formally verified by the Lean kernel into a reusable lemma pool. Evaluated on six designs, the method successfully produced 403 theorems, of which 287 foundational lemmas were automatically reusable, achieving a reuse rate of 80.2%. This significantly enhances the automation and scalability of formal RTL verification.