BEHAVE: Functional Behavior Modeling Enables Self-Improving Agents for Hardware Design and Verification
This study addresses the issue of false rejections of valid hardware designs caused by cycle-level matching in design verification. To overcome this, we propose BEAVE, a framework that enables multi-round co-design and verification through behavioral models. The core innovation lies in introducing a Behavior Intermediate Representation (IR) to decouple functionality from timing, thereby constructing a verifiable reinforcement learning reward mechanism without requiring reference RTL. Furthermore, by integrating random sampling with solver-guided search, the framework supports PPA exploration and self-improvement. Experimental results demonstrate that this approach increases the RTL pass@1 rate of Qwen3.8-27B from 55% to 75%, achieving performance comparable to reinforcement learning conducted on large-scale task pools.