Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear reliability of System 1 models in automated safety decision-making and the issue that aggregated metrics obscure failures against specific attacks. We evaluate the capability of models such as Jev to detect prompt injection risks through a multi-model comparative assessment integrating accuracy, probability calibration, and selective automation strategies. Our findings reveal hidden failure modes within specific attack groups despite strong overall performance, demonstrating that single confirmation mechanisms cannot guarantee valid error bounds on test sets and that automated approval rates remain extremely low under strict constraints. The primary contribution of this work lies in establishing a joint evaluation paradigm that integrates accuracy, calibration quality, and automated decision-making effectiveness for assessing model-based safety systems.
📝 Abstract
Model-based judges support agent security by detecting prompt injections, assessing interaction risks, and screening harmful requests. System One models expose typed decisions with probabilities that software can use to allow, block, or escalate inputs, but whether these probabilities support reliable automated security decisions remains unclear. We evaluate Jev, Laya, Decider, and Bespoke Nimble against specialized classifiers and language-model judges, examining decision accuracy, probability calibration, and selective automation. We identify three main findings. (1) Strong overall performance and favorable average calibration can hide systematic failures on particular attack groups, including attacks that models confidently classify as safe. (2) Under strict limits on missed attacks, the evaluated policies allow few inputs automatically, and choosing separate allow and block thresholds increases automation mainly by blocking more inputs. Even policies that meet error limits during validation can exceed them on unseen inputs. (3) A second judge can detect some missed attacks, but it may also reject more benign inputs and repeat the first model's high-confidence errors. These findings show that model accuracy, probability calibration, and the behavior of the resulting decision policy must be evaluated together.
Innovation

Methods, ideas, or system contributions that make the work stand out.

System One models
Agent security
Probability calibration
Selective automation
Prompt injection