PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck in real-time proactive safety defense for LLM agents, which stems from the scarcity of causally consistent data during environment interactions. To this end, we propose a high-fidelity trajectory synthesis framework. Methodologically, we design a progressive trajectory unfolding mechanism coupled with reasoning-enhanced causal correction to mitigate safety drift. Furthermore, multi-model arbitration annotation, bilingual benchmark construction, and cultural localization strategies are introduced to enforce causal consistency constraints. Experimental results demonstrate that the proposed approach achieves an F1 score of 91.46% on unsafe instances and reduces the attack success rate from 20.82% to 0.40%, thereby establishing efficient real-time safety guardrails for LLM agents.
📝 Abstract
The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PROACT-Agent, a framework for synthesizing high-fidelity trajectories to enable real-time guardrails. We identify a critical "safety drift" in prior benchmarks, where lenient annotation paradigms fail to enforce temporal consistency. PROACT-Agent addresses this through: (1) Progressive Trajectory Unrolling to reveal risks hidden in long-context interactions; (2) Reasoning-Augmented Causal Rectification to enforce monotonic causal consistency; and (3) Culturally-Aware Data Localization for cross-border robustness. We introduce PROACT-Bench, a bilingual safety benchmark with 155,780 states labeled through multi-model adjudication. Evaluating updated context before the next LLM inference, the trained guard achieves 91.46% unsafe-class F1 and 90.63% exact-boundary detection under complete source holdout. In AgentDojo, it reduces non-DoS targeted attack success from 20.82% to 0.40%.
Problem

Research questions and friction points this paper is trying to address.

LLM Agents Safety
Runtime Intervention
Safety Drift
Causal Consistency
Real-time Guardrails
Innovation

Methods, ideas, or system contributions that make the work stand out.

Progressive Trajectory Unrolling
Causal Rectification
Real-time Guardrails
Safety Drift
PROACT-Bench
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Ding Jia
State Key Laboratory of Intelligent Vehicle Safety Technology, Changan Automobile
W
Wei Liu
State Key Laboratory of Intelligent Vehicle Safety Technology, Changan Automobile
X
Xianglong Du
State Key Laboratory of Intelligent Vehicle Safety Technology, Changan Automobile
Y
Yingjie Li
State Key Laboratory of Intelligent Vehicle Safety Technology, Changan Automobile
Y
Yingqing Yang
State Key Laboratory of Intelligent Vehicle Safety Technology, Changan Automobile
H
Huili Yu
State Key Laboratory of Intelligent Vehicle Safety Technology, Changan Automobile
Z
Zhangsong Zhan
State Key Laboratory of Intelligent Vehicle Safety Technology, Changan Automobile
Chu Zhou
Chu Zhou
National Institute of Informatics
Computational Photography