🤖 AI Summary
This study addresses the inability of LLM agents to autonomously report errors due to the lack of external verification, which typically forces reliance on budget exhaustion for termination. Drawing inspiration from the V-model in software engineering, this work proposes a hierarchical verification architecture. Its core innovation lies in decoupling zero-cost deterministic gating from optional LLM-based judging, enforced by a controller to achieve precise fault localization, proactive halting, and graceful degradation. Experiments on multi-hop question answering using the MuSiQue dataset demonstrate that, with an 8B-parameter model, the proposed architecture successfully answers eight out of ten questions while correctly refusing unanswerable ones, whereas the baseline fails entirely. Notably, most corrections are accomplished through the zero-cost gates, significantly enhancing overall system reliability.
📝 Abstract
Large language model (LLM) agents built on the ReAct pattern concentrate four responsibilities in one model: selecting a strategy, choosing each action, formatting it, and judging whether the result is adequate. Nothing outside the generative loop can reject its output, so an agent that cannot make progress does not report failure; it runs until an external budget stops it. We propose treating verification as an architectural layer by adapting the V-model from software engineering: specification levels descend from requirements to individual steps, each level is paired with a dedicated verifier, a deterministic controller enforces every verdict, and only verification outcomes write to memory, so a rejection localizes the level that introduced the fault and an agent halts by declining rather than by exhaustion. Each verifier separates a zero-cost deterministic \emph{gate} from an optional LLM \emph{judge}, so the contribution and cost of each can be measured independently. We report a pilot implementing the acceptance- and unit-level verifier pairs, comparing five configurations that share one executor, tool set, and scorer and differ only in verification, on the four-hop stratum of MuSiQue with an 8B-parameter backbone. Across 47 executions, the two unverified configurations answered none of ten questions, every run ending at a step cap or provider token limit; the verified configuration without a planner answered eight and abstained on the rest. Deterministic gates produced eight of the nine observed corrections at zero marginal cost, and planning degraded performance once verification was present. These results characterize termination behavior, not accuracy at scale; we outline a twelve-month plan to complete and evaluate the full architecture, including the integration-level pair the pilot omits.