🤖 AI Summary
This study addresses the challenge of verifying defects in AI research—particularly errors relying on prior knowledge—when only outputs are available. To this end, it introduces a novel "research contract" mechanism that binds experimental choices, execution obligations, and evidence, thereby establishing verification boundaries under information asymmetry. The proposed approach enables automated verification through deterministic checkers, registration of fault-specific rules, and metadata filtering, while explicitly distinguishing contract-relative verification from scientific truth. Experimental results demonstrate that the system successfully detected all eight registered mutations and identified 104 defects across 144 variant cases, effectively ensuring the reliability of compliance verification.
📝 Abstract
Some defects in an AI-generated study can be identified from its artifacts; others require knowledge of what was approved before execution. We propose study contracts that bind declared experimental choices, run obligations and claim scope to recorded execution evidence, and distinguish this contract-relative verification from scientific truth. A diagnostic using eight self-authored clean/mutated pairs illustrates the information boundary. A deterministic checker applying a registered, fault-specific rule to approved and executed objects detected all eight registered mutations. Across eighteen recorded judge aliases given individual metadata-filtered packages without pair context or the registry-selected fault label, 104 of 144 mutated evaluation cases received defect flags; the remaining cases comprised 32 abstentions and eight terminal failures, with no explicit clean decisions on mutated cases. Some packages retained approval and execution fields, including digests. The prompt instructed judges to abstain when evidence was insufficient. These results characterize a deliberately information-asymmetric development setting; they do not isolate the effect of authoritative information from differences in task specification and rule selection, and they are not comparative verifier quality or agent reward hacking. We identify full-information comparisons, legitimate-adaptation controls and closed-loop agent evaluations as necessary tests of whether contract checks improve useful compliant completion under optimization.