🤖 AI Summary
This study addresses a critical security vulnerability in existing agent safeguard models, which focus exclusively on prohibited actions while neglecting unfulfilled safety obligations. To transcend traditional safeguarding paradigms, this work is the first to formally define and quantify the risk of unfulfilled obligations. Methodologically, we construct ObligationBench, the first benchmark for obligation identification, and train the ObligationGuard model using 40,000 synthetic data samples integrated with expert-validated trajectories to enable fine-grained obligation recognition and evaluation. Experimental results demonstrate that ObligationGuard significantly outperforms existing baselines in both recall and exact match rates. By effectively mitigating these previously overlooked safeguarding blind spots, the proposed approach substantially enhances the overall security capabilities of AI agents.
📝 Abstract
Guard models are increasingly used to safeguard LLM-based agents, primarily by identifying actions that agents are forbidden to perform. However, identifying forbidden actions alone is insufficient to ensure agent safety. In this paper, we argue that agent safety also depends on identifying required yet unperformed safety-critical actions, which we call obligations. Our preliminary study on a popular benchmark for evaluating safety shows that 56.92% of GLM-5.3 trajectories contain unfulfilled obligations, compared with only 30.00% containing forbidden actions. This finding reveals unfulfilled obligations as a major and previously overlooked source of safety risk. However, to our knowledge, no existing benchmark evaluates whether guard models can identify these obligations. To close this gap, we introduce ObligationBench, the first benchmark for evaluating the capability of obligation identification, comprising 240 expert-validated trajectories covering issue resolution, feature development, and terminal operations. Our evaluation of 14 representative models reveals substantial limitations: the highest recall and exact-match rate are only 48.97% and 10.00%, respectively. To address these limitations, we develop ObligationGuard using 40,000 synthetic training examples. ObligationGuard achieves 57.52% recall and an exact-match rate of 21.67%, surpassing all evaluated models on both metrics. We call on the community to incorporate obligation identification into the design and evaluation of future guard models to improve agent safety.