🤖 AI Summary
This study addresses cross-object tampering and audit failures in tool-use agents caused by the separation of declarations, evidence, and execution logs. We propose a declaration-anchored execution contract mechanism that jointly binds declarations, source snippets, and execution prefixes to enable replayable integrity auditing. A core innovation is the first architecture decoupling structural from semantic support, defining seven independently testable properties to ensure evidence immutability while integrating domain-isolated commitments, hash verification, deterministic validation, and conflict-aware guards. Experimental results demonstrate that this approach achieves a 99.61% detection rate across 1,280 attacks, with a guard F1-score of 0.8865 and a false acceptance rate of only 7.29%.
📝 Abstract
Tool-using agents can expose citations and execution logs while leaving a critical association unaudited: whether the claim shown to a user is the claim emitted by the committed execution and supported by the cited source. A valid citation and a valid trace can therefore remain individually well formed while being transplanted across claims, actions, runs, or source versions. We introduce a claim-anchored execution contract that jointly binds the emitted claim, its exact source span, the ordered execution prefix that produced it, and the source version and access state observed by that execution. Each receipt contains an emission anchor that deterministically locates the claim inside a committed answer or claim-bearing action, together with source identifiers, offsets, hashes, quotes, and a domain-separated execution commitment. A deterministic integrity verifier reconstructs these bindings before semantic or task labels are joined. We separate this integrity plane from a pluggable support plane, so structural validity is not used as a proxy for entailment. The contract exposes seven independently testable properties: claim-emission binding, source binding, ordered-execution binding, oracle separation, persisted-object replay, execution-rerun consistency, and version/access binding. Across 1,280 cross-object attacks, the joint contract detects 1,275 substitutions (0.9961). Removing a targeted property reduces its attack-detection rate to 0.0156-0.0625. On an independently adjudicated 384-pair split, the conflict-aware support guard reaches F1 0.8865 and false acceptance 0.0729; on unseen failure families, these rates are 0.8679 and 0.0938.