GRADE-RTL: Evaluating LLM-Generated RTL Beyond Compilation

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出GRADE-RTL框架,通过五项检查评估大语言模型生成的RTL代码的结构完整性、功能正确性和实现效率,超越编译层面。
📝 Abstract
Large language models (LLMs) can generate register-transfer-level (RTL) code from natural-language specifications, but compilation alone does not establish structural completeness, functional correctness, or implementation efficiency. This paper presents a framework for evaluating LLM-generated RTL beyond compilation, which we named GRADE-RTL. Using fixed prompts and validation settings, GRADE-RTL applies five checks, such as Port Signature, Compilation, Elaboration, Module Completeness, and Functional Equivalence against trusted reference RTL. With the help of compact metrics, we identify where candidates fail and distinguish end-to-end success from eligibility for downstream synthesis. We evaluate nine general-purpose and RTL-specialized LLMs on ten edge-relevant intellectual property designs under a budget of three generation attempts with failure-directed feedback. End-to-end success ranges from 0% to 70% across the evaluated models, with failures extending beyond compilation to hierarchy resolution, incomplete logic, and behavioral mismatch. FPGA implementation and 65 nm ASIC synthesis results further show that functionally equivalent RTL can differ substantially in resource use, timing, area, and power. A PID-controller place-and-route case study illustrates the physical-design consequences of different RTL implementations. GRADE-RTL provides a practical basis for comparing LLM-generated hardware descriptions by separating structural validity, behavioral correctness, and implementation quality.
Problem

Research questions and friction points this paper is trying to address.

Large language models
Register-transfer-level code
Structural completeness
Functional correctness
Implementation efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

GRADE-RTL
LLM-generated RTL
Functional Equivalence
Implementation Quality
End-to-end Success
🔎 Similar Papers
No similar papers found.