🤖 AI Summary
This study addresses the inability of large language models (LLMs) to reliably determine rule-checking boundaries, which exposes AI verifiers in regulatory scenarios to severe misjudgment risks. We provide the first formal proof that LLMs lack self-determination capabilities and propose CoVer, a framework incorporating a “corroborate-then-verify” mechanism that discards the flawed assumption of model consensus. By integrating multi-model cross-validation, synthetic intervention testing, and deterministic paraphrasing evaluation, CoVer constructs reliable verification boundaries. Experimental results demonstrate that reference judges fail to eliminate blind spots, whereas CoVer significantly reduces false positive rates across six corpora. These findings establish the necessity of constructive verification boundaries for dependable AI regulation.
📝 Abstract
A verifier for an agent faces rules of two kinds: the ones a fixed check can settle and the ones that require a judge. A team that derives its own checks fixes that split up front. Where the requirements come from outside, as in finance, healthcare and law, the agent enforces rules it did not write, so the split falls to runtime, recurring for every predicate of every rule on every action at a rate no reviewer can audit. Every escalation scheme assumes a model can make that decision itself, that it is self-decidable. Across six corpora, including the EU AI Act, FINRA guidance and a deployed credit agent, we collect roughly 22,000 labels from four models built by three labs. They agree almost perfectly where the answer is obvious and collapse on regulatory text; their errors run in opposite directions, so no model can be trusted as the conservative choice; and on the deployed agent's own rule-set they err together, over-claiming that a fixed check will do, the direction that never gets escalated. We introduce CoVer (corroborate-then-verify), which treats unanimity as a nomination, admitting a predicate only when the check synthesized for it survives intervention, reading fields the agent cannot write and holding under deterministic rewording. That gate rejects most of what corroboration wrongly admits, at a cost in coverage we report rather than tune away. The obvious alternative, agreement with a reference judge, certifies nothing: it climbs from 30% to 77% across calibration bands while the genuinely decidable share does not move, because a judge drawn from the population under indictment ratifies the blind spot it shares. Self-decidability is not a capability to elicit from a model but a boundary the verifier must construct.