AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of disentangling over-refusal from task failure and the difficulty of safety evaluation in tool-use agents by proposing AgentBound. This framework introduces a novel four-way counterfactual generation and evaluation system that leverages trajectory-level and post-state decision workflow transformation techniques to independently manipulate risk and permission variables, effectively decoupling surface risk, action permissibility, and task capability. Furthermore, it incorporates a lightweight runtime calibration module for optimization. Experiments across 17 configurations demonstrate that AgentBound improves authorized task completion rates by an average of 18.2% while increasing unsafe action interception rates by 5.4%, thereby achieving synergistic enhancements in both safety and usability.
📝 Abstract
Safety alignment for large language models (LLMs) in conversational settings is largely framed around whether to answer or refuse a request. In agentic settings, however, the same models must decide whether to act as permission-critical evidence emerges during execution. This creates a distinct challenge: apparent risk, action permissibility, and task competence are easily confounded, making agentic over-refusal difficult to distinguish from ordinary task failure. To address this, we introduce AgentBound, the first four-way counterfactual generation-and-evaluation framework for tool-using agent safety. AgentBound transforms the same executable workflow by independently varying apparent risk and action permissibility, enabling controlled comparisons of risky-looking but authorized tasks and routine-looking but unauthorized tasks. These comparisons jointly diagnose over-refusal and unsafe compliance while controlling for task competence. We instantiate AgentBound as a human-validated 4,000-task evaluation suite with trajectory-based and post-state-based judgments. Across 17 model and harness configurations, high safety frequently coexists with poor authorized-task completion: GPT-5.5 blocks 99.5\% of routine-looking unauthorized actions yet completes only 28.7\% of risky-looking authorized tasks. We further train a lightweight runtime calibration module that improves authorized-task completion by 18.2\% on average across 10 evaluated configurations, while improving unsafe-action blocking by 5.4\% on average. These show that effective agentic alignment requires action decisions to track permission-relevant execution evidence, rather than refusal strength alone.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
safety alignment
over-refusal
action permissibility
counterfactual evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Evaluation
Agentic Safety
Tool-Using LLM Agents
Over-refusal
Runtime Calibration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.