Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification

📅 2026-04-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inefficiency of coverage convergence in hardware verification, where existing large language model (LLM) agents lack systematic analysis of hard-to-reach coverage holes and principled mechanisms for allocating reasoning resources. The authors propose a two-tier agent framework that integrates a base Codex agent with a domain-enhanced LangGraph system, establishing the first taxonomy of coverage gaps—categorized by methodological ceilings and reasoning frontiers—to expose fundamental limitations of purely LLM-driven verification. They introduce a profile-driven agent design paradigm, incorporating fine-grained tracking of token consumption across six categories and a coverage feedback loop. Evaluated on multiple designs, the approach achieves 95–99% coverage while reducing token usage by 4–13× and accelerating convergence by 2–4× compared to general-purpose baselines.

Technology Category

Cognitive Modeling & Cognitive Systems: Agent ArchitecturesMultiagent Systems: Agent CommunicationMachine Learning: Large Multimodal Models (LMMs)

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Coverage closure is the most time-consuming phase of hardware verification, and recent large language model (LLM)-based coding agents offer a promising approach to automated stimulus generation. However, prior LLM-based flows do not systematically analyze which coverage holes remain difficult to close or how inference-time computation is allocated during agentic verification. As a result, the efficiency limits and failure modes of LLM-based coverage closure remain poorly understood, particularly for large designs. We present an empirical study using a two-tier agentic framework comprising a base Codex agent and an enhanced domain-specialized LangGraph system. Our framework enables a taxonomy of coverage holes: methodology-bound ceilings (integration tied-off hardware, infeasible boundaries, dead code) and reasoning frontiers (protocol sequencing, multi-module pipeline warm-up, narrow timing conditions), exposing fundamental limits of purely LLM-driven approaches. We further instrument the system to track token usage across six categories, including system prompt, design comprehension, stimulus generation, coverage feedback, error recovery, and agentic overhead. We show that domain specialization shifts token allocation toward coverage-directed reasoning and improves efficiency. Across designs, the enhanced system achieves comparable or higher coverage (95-99%) while using 4-13x fewer tokens and converging to coverage targets 2-4x faster than a general-purpose baseline. Our results characterize the limits of LLM-based coverage closure, inform benchmark design and human escalation strategies, and guide profile-driven agent design for hardware verification.
Problem

Research questions and friction points this paper is trying to address.

coverage closure
token allocation
hardware verification
LLM-based agents
coverage holes
Innovation

Methods, ideas, or system contributions that make the work stand out.

coverage closure
LLM-based verification
token allocation
domain-specialized agents
hardware verification
V
Vihaan Patel
Arizona State University
V
Vidya Chhabria
Arizona State University
A
Aman Arora
Arizona State University