🤖 AI Summary
This study addresses the issue of false rejections of valid hardware designs caused by cycle-level matching in design verification. To overcome this, we propose BEAVE, a framework that enables multi-round co-design and verification through behavioral models. The core innovation lies in introducing a Behavior Intermediate Representation (IR) to decouple functionality from timing, thereby constructing a verifiable reinforcement learning reward mechanism without requiring reference RTL. Furthermore, by integrating random sampling with solver-guided search, the framework supports PPA exploration and self-improvement. Experimental results demonstrate that this approach increases the RTL pass@1 rate of Qwen3.8-27B from 55% to 75%, achieving performance comparable to reinforcement learning conducted on large-scale task pools.
📝 Abstract
Developing agents for hardware design and verification requires reliable correctness feedback. As a hardware specification may permit correct implementations with different latencies, matching design and reference outputs cycle by cycle can reject valid designs. To address this, we introduce BEHAVE, an agentic framework for multi-turn joint hardware design and verification through functional behavior modeling. We define Behavior IR to express task functionality as executable behavior models without prescribing implementation timing beyond the specification. The agent iteratively develops a register-transfer-level (RTL) design and a behavior model as the design's verification reference. Our evaluator, BEHAVE-Sim, checks both artifacts separately against a hidden golden behavior model using input stimuli generated by random sampling and solver-guided search. BEHAVE thus supports power, performance, and area (PPA) exploration across task-permitted latencies and microarchitectures. During training, the same evaluator provides verifiable reinforcement learning (RL) rewards from specification-behavior pairs without reference RTL. For self-improvement, the agent continually searches for high-level implementations relevant to its capability gaps, constructs and checks specification-behavior pairs, and trains on the expanded task pool. We release BEHAVE-Train and BEHAVE-Eval with 600 human-reviewed specification-behavior pairs for realistic hardware workloads. Starting from 60 seed tasks and acquiring 100 new tasks, self-improvement raises Qwen3.8-27B's RTL pass@1 on BEHAVE-Eval from 55.0% to 75.0%, reaching performance comparable to RL using a 540-task pool.