The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents

๐Ÿ“… 2026-08-04
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge of verifying long-horizon agents operating under untrusted self-reported states and narratives. It proposes a structured self-verification architecture in which a deterministic execution module governs all beliefs, while a language model may only submit typed proposals; such proposals are accepted only if their pre-registered predictions align with subsequent observations as verified by code-level comparison. The approach introduces a novel self-nullifying verification mechanism and an invisible shadow reference system, enabling, for the first time, decoupled measurement of commitment drift and binding driftโ€”even in regions lacking explicit mechanisms, where drift metrics remain well-defined. Ablation studies show that removing the commitment mechanism raises goal abandonment to 1.00 while binding errors stay at zero, and omitting binding repair induces no stepwise drift but triggers upstream assumption collapse. Although none of the 52 runs completed the task, the efficacy of the proposed verification methodology is robustly demonstrated.
๐Ÿ“ Abstract
How do you verify a long-horizon agent when its own state and self-reports are exactly what you cannot trust? We present an agent instrument built so that verification is structural rather than post-hoc. A deterministic Executive owns all belief; a language model may only file typed proposals, and a claim is admitted only when a prediction pre-registered before acting is matched against observation by code. Two properties make the instrument a verifier of its own science, not just of the agent: every run invalidates itself when per-organ write-error, render-size, or salted-canary-echo floors are breached (four of the first eight architecture runs were invalidated, each localizing a real defect); and a render-invisible shadow reference compiles the plan the full system would have committed in every ablation cell, so drift metrics are defined even where the mechanism under test has been removed. Using this instrument we report a clean, single-variable result on a failure every long-horizon agent suffers: ablating the commitment mechanism flips goal-abandonment from 0.00 to 1.00 while binding error stays flat at 0.00 (three seeds per cell, up to 394 reference beats per run, every run gated valid). The binding channel, by contrast, does not reappear as per-beat drift when its repair is ablated -- because binding is code-owned, the failure class is structurally absorbed, its only residue appearing one layer upstream as a collapse in hypothesis formation. We report these under full disclosure that task efficacy is null (zero level completions across 52 gated runs on ARC-AGI-3), pre-registered as a structural defeater. The contribution is a verification methodology for agent development and the drift decomposition it makes measurable.
Problem

Research questions and friction points this paper is trying to address.

long-horizon agents
verification
commitment drift
binding drift
self-reporting untrustworthiness
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-verifying agent
commitment drift
binding drift
structural verification
ablation shadow reference