Efficient Auditing of Adversarial AI Agent Behavior from Agent Traces

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of static rules to evasion and the prohibitive cost of exhaustive auditing in monitoring large language model (LLM) agent behaviors. We propose a two-stage adaptive framework that synergizes rule-based filtering with LLM-driven analysis. The framework first performs efficient preliminary screening using single-event and sequential rules, followed by precise context-aware auditing executed by an LLM, with both rules and audit instructions jointly optimized based on empirical data. Experimental evaluations on the OpenAgentSafety benchmark demonstrate that our approach reduces audit frequency by 72% and token consumption by 70%. Furthermore, in multi-agent scenarios, it achieves complete interception of malicious behaviors while decreasing token overhead by over 80%, effectively balancing security assurance with computational efficiency.
📝 Abstract
AI agents powered by large language models (LLMs) can perform complex tasks but may harm the systems they operate in, either intentionally or unintentionally. Existing agent monitoring approaches rely on rule-based guardrails or LLM-based trace auditing. However, rule-based guardrails can be bypassed through obfuscation and may miss harmful actions beyond their predefined rules, whereas applying an LLM to audit every action is costly. We present a two-stage agent trace auditing framework. The first stage uses single-event and trace-sequence rules to select pending actions for inspection; the second uses an LLM audit agent to examine each selected action in the context of the agent's preceding trace before execution. We jointly refine the gate rules and audit instructions using training data, allowing the framework to adapt to complex agent behaviors rather than relying solely on predefined rules. On the public benchmark OpenAgentSafety, our framework reduces the average number of LLM audits from 8.15 to 2.33 per run and token usage from 47.8k to 14.6k, with a detection rate of 72.8\% compared with 81.5\% when every action is audited. In two simulated multi-agent case studies, the framework flags all malicious traces while reducing audit token usage by more than 80\%.
Problem

Research questions and friction points this paper is trying to address.

AI agent safety
adversarial behavior auditing
LLM-based agents
trace monitoring
audit efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Two-stage auditing framework
Agent trace auditing
Joint optimization
Adversarial AI agents
Cost-efficient monitoring
🔎 Similar Papers
No similar papers found.