Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

📅 2026-07-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of reproducibility in current agent safety evaluations, where identical surface-level outcomes may stem from vastly different evidentiary bases, undermining validity verification. To resolve this, the authors propose a cross-platform, vendor-neutral reproducibility measurement framework that introduces, for the first time, decision-level reproducibility metrics and a claim-evidence overstatement gap. The approach incorporates an evidence sufficiency card and a release gating mechanism, built upon a counterfactual replay intervention protocol, replay precondition probes, cross-framework adapters, and a twelve-dimensional evidence scoring system. Crucially, it enables evaluation on both public and bundled trajectories without requiring new model executions. Experiments demonstrate that semantically equivalent inputs yield sufficiency scores ranging from 0.458 to 0.833; the release gate successfully blocks low-scoring variants (0.542) while approving high-scoring versions (0.667). All results are fully reproducible via an open-source package.
📝 Abstract
Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may sit atop materially different evidence regimes. No vendor-neutral, runnable instrument scores reconstructability as an evaluation-validity metric: whether captured evidence can reconstruct the decision a claim depends on. This paper introduces a property-level reconstructability metric over eight decision-property classes and a cross-harness adapter emitting per-decision Evidence Sufficiency Cards backing a per-run monitor-coverage release check. It specifies a counterfactual-replay intervention protocol, implements its replayability-precondition probe, and defines a claim-evidence overclaim gap. On public and bundled traces, without new model runs, twelve-field sufficiency spans 0.458-0.833 across four inputs sharing a surface reading; replay preconditions are unmet in every scored trace. In a synthetic release-gate pair, the sufficiency gate blocks the raw variant (0.542) and passes the instrumented (0.667). Safety-evaluation claims should travel with their reconstructability vector; a reproducibility package regenerates every reported number.
Problem

Research questions and friction points this paper is trying to address.

agent-safety
reconstructability
evaluation validity
evidence sufficiency
decision reconstruction
Innovation

Methods, ideas, or system contributions that make the work stand out.

reconstructability
evidence sufficiency
counterfactual replay
agent safety evaluation
cross-harness adapter
🔎 Similar Papers
2024-08-192024 IEEE 27th International Conference on Intelligent Transportation Systems (ITSC)Citations: 0