🤖 AI Summary
This study investigates the selective retraction capability of language agents when evidence changes or instructions are revoked, specifically their ability to precisely suspend affected actions while preserving valid work. We construct a synthetic environment, NAQD-Env, alongside a selective retraction benchmark grounded in deterministic reference policies to jointly evaluate policy consistency, task value, and recovery performance. Experiments reveal that existing models exhibit extremely low retraction recall and lack recovery capabilities. While exploratory fine-tuning significantly improves decision accuracy, it induces over-retraction and the loss of event reports. This work exposes critical deficiencies in the dynamic adaptability of current agents and offers new directions for designing trustworthy AI systems.
📝 Abstract
Language agents must revise planned actions when evidence changes, permission is revoked, or a stop instruction arrives. A useful response is selective: suspend affected actions, preserve unaffected work, and resume only after sufficient repair. We introduce NAQD-Env, a synthetic environment that evaluates these decisions against a deterministic reference policy over explicit evidence, authorization, and constraint dependencies. Eleven dependency families support evaluation on development structures, held-out families, and held-out combinations of structures. Metrics distinguish attempted violations from violations permitted by a simulated execution gate and jointly report policy agreement, task value, withdrawal, resumption, and event reporting. We evaluate three open-weight instruction-tuned models from two families under three prompt conditions on 350 frozen scenarios, yielding 3,150 model-prompt episodes before gate replay. Across the reported conditions, withdrawal recall is at most 0.06, no valid resumption is observed at eligible opportunities, and only one episode matches the complete reference policy. Under the NAQD prompt, Qwen2.5-7B has fewer unsafe-attempt episodes than Qwen2.5-3B and Llama-3.1-8B, but also completes less useful work and preserves unaffected actions less accurately. Exploratory supervised fine-tuning probes increase Qwen2.5-3B decision accuracy from 0.45-0.54 to 0.83-0.92; separate diagnostics reveal inappropriate withdrawal after curriculum omissions and a loss of event reporting. These results motivate evaluating selective withdrawal as a distinct component of agent reliability. The setting measures policy application with trusted structured inputs and does not establish real-world containment or source-verification ability.