🤖 AI Summary
This study addresses the lack of reproducibility in current agent safety evaluations, where identical surface-level outcomes may stem from vastly different evidentiary bases, undermining validity verification. To resolve this, the authors propose a cross-platform, vendor-neutral reproducibility measurement framework that introduces, for the first time, decision-level reproducibility metrics and a claim-evidence overstatement gap. The approach incorporates an evidence sufficiency card and a release gating mechanism, built upon a counterfactual replay intervention protocol, replay precondition probes, cross-framework adapters, and a twelve-dimensional evidence scoring system. Crucially, it enables evaluation on both public and bundled trajectories without requiring new model executions. Experiments demonstrate that semantically equivalent inputs yield sufficiency scores ranging from 0.458 to 0.833; the release gate successfully blocks low-scoring variants (0.542) while approving high-scoring versions (0.667). All results are fully reproducible via an open-source package.
📝 Abstract
Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may sit atop materially different evidence regimes. No vendor-neutral, runnable instrument scores reconstructability as an evaluation-validity metric: whether captured evidence can reconstruct the decision a claim depends on. This paper introduces a property-level reconstructability metric over eight decision-property classes and a cross-harness adapter emitting per-decision Evidence Sufficiency Cards backing a per-run monitor-coverage release check. It specifies a counterfactual-replay intervention protocol, implements its replayability-precondition probe, and defines a claim-evidence overclaim gap. On public and bundled traces, without new model runs, twelve-field sufficiency spans 0.458-0.833 across four inputs sharing a surface reading; replay preconditions are unmet in every scored trace. In a synthetic release-gate pair, the sufficiency gate blocks the raw variant (0.542) and passes the instrumented (0.667). Safety-evaluation claims should travel with their reconstructability vector; a reproducibility package regenerates every reported number.