AgentSpy: Making AI Agent Behavior Observable

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of unobservable internal behaviors, incomplete trajectory logging, and latent security risks in AI agents by proposing a dual analysis framework for compliance and safety based on external observation. Methodologically, agents are executed within sandboxed environments where system calls and network traffic are captured via declarative configurations. By integrating a deterministic rule engine, this approach overcomes the limitations of relying solely on agent self-reported trajectories, enabling comprehensive monitoring of both primary processes and their subprocesses. Experimental results demonstrate that the proposed method achieves a 92.2% similarity across repeated executions, identifies 18% of extraneous activities, and successfully detects four out of five attack categories with zero false positives.
📝 Abstract
AI agents built on large language models (LLMs) run shell commands, read and write files, and reach the network, typically with their user's privileges. However, what an agent does during an execution is difficult to understand: tests assert on the result, and the agent's trajectory records only what the agent reports about itself, which may omit behavior executed by its subprocesses. We present AgentSpy, an approach that observes an agent from outside the agent. AgentSpy runs the agent in an isolated environment, configured by a declarative specification, and records the system calls and network traffic of the agent and of every process it executes. Based on this monitoring, AgentSpy supports two families of analyses: conformance analyses, which measure obligations, i.e., what an agent execution should do, and safety analyses, which check prohibitions, i.e., what an agent execution must never do. We instantiate one analysis of each family. The reliability analysis uses rules to summarize each run by the environment resources the agent uses: the commands it executed, the files it accessed, and the hosts it contacted. The security analysis applies deterministic rules to the system calls of an execution. For reliability, we evaluated AgentSpy on 77 tasks with the codex harness and three recent LLMs, executing each task three times. Sets of repeated runs of the same task are more similar than sets that include runs of another task in 92.2% of the comparisons. Among tasks for which all three runs pass outcome-based tests, the agent performs task-unrelated activities in 18% of the cases, reads the grading files in 7%, and does not use the developers'guidance in 17%. For security, the generic rules of AgentSpy detect four of five attack categories we considered, with no false positives across 50 runs.
Problem

Research questions and friction points this paper is trying to address.

AI agent observability
LLM-based agents
behavior monitoring
safety analysis
conformance analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI agent observability
system call monitoring
declarative specification
conformance analysis
safety analysis