🤖 AI Summary
This study addresses the absence of evaluations for industrial-grade VHDL and repository-level, multi-file generation in existing large language model (LLM) benchmarks by introducing the first large-scale VHDL repository-level benchmark. Encompassing over one hundred open-source libraries with self-verifying testbenches, the proposed benchmark employs structured problem definitions and module stubbing techniques. Coupled with multi-step reasoning strategies such as Reflexion, it systematically evaluates LLM performance across syntax, semantics, and cross-file reasoning. This work bridges a critical gap in hardware description language evaluation by focusing on VHDL, revealing significant limitations of current models in multi-file coordination and hierarchical design comprehension. Ultimately, this research provides the community with an essential evaluation resource for advancing LLM capabilities in complex hardware design tasks.
📝 Abstract
Large Language Models (LLMs) are increasingly applied in hardware design automation, demonstrating strong potential in generating and understanding hardware description languages. However, most existing benchmarks focus on Verilog, with limited evaluation of VHDL, which remains widely used in industry and academia for FPGA and safety-critical systems. To address this gap, we introduce VHDL-REPOBENCH, a large-scale, cross-file, repository-level benchmark for assessing LLM capabilities on realistic VHDL design generation and analysis tasks. VHDL-REPOBENCH curates ~100 open-source VHDL repositories, encompassing ~2.5k VHDL files and ~500 testbenches, and provides structured problem statements, module stubs, and self-verifying testbenches. The benchmark enables comprehensive evaluation across syntax, semantic correctness, hierarchical reasoning, cross-file dependency resolution, and functional verification. We evaluate several state-of-the-art models, including GPT-4o, Llama-3-70B, Qwen2.5-72B, CodeLlama-70B, and multi-step reasoning approaches such as Reflexion and CoDes. Results reveal that while current LLMs achieve moderate line- and block-level accuracy, substantial challenges remain in multi-file reasoning, hierarchical design understanding, and specification-to-module generation. VHDL-REPOBENCH represents the first large-scale VHDL-focused benchmark and provides a valuable resource for the hardware design community to evaluate, compare, and advance LLM capabilities for practical VHDL development.