🤖 AI Summary
This study addresses the lack of joint task- and stage-level annotations in existing traffic datasets, which hinders privacy-preserving auditing of large language model (LLM) agent behaviors. To this end, we construct ANT, a multi-granularity agent network traffic dataset that provides the first three-tier fine-grained semantic annotations encompassing risks, scenarios, and behavioral primitives. By integrating network traffic analysis with agent execution monitoring, ANT establishes a unified benchmark evaluation framework. Experimental results reveal significant limitations of existing methods in distinguishing similar traffic patterns and identifying malicious workflows. This work bridges the critical data gap in agent traffic auditing, laying the foundation for precise, non-intrusive supervision of LLM agent behaviors.
📝 Abstract
The growing adoption of large language model (LLM) agents creates a need for network administrators and security teams to audit agent behavior within organizational networks without inspecting private user content. Network traffic offers an observable source of evidence, but how much it reveals about agent tasks and operations remains unclear. Existing traffic datasets lack the joint task and stage annotations needed to evaluate this question. We introduce ANT (Agent Network Traffic), a dataset providing agent behavior information at risk, scenario, and behavior primitive granularities alongside network traffic. ANT contains 3,114 execution episodes across 20 tasks and five scenarios, comprising 276,417 bidirectional flows and 40,049 behavior primitive segments organized into 47 macro groups. We establish a benchmark for agent risk identification, scenario recognition, and behavior primitive classification using 13 representative traffic analysis baselines. The results show that existing methods recover useful but uneven behavioral signals. They struggle to identify risk when malicious workflows resemble benign tasks and to distinguish scenarios with similar traffic patterns. Primitive classification is more reliable for frequent macro groups and those with distinctive traffic patterns than for rare or semantically similar groups. ANT provides a common basis for developing more precise auditing and forensic analysis of agent behavior from network traffic. Our data and code are available at https://anonymous.4open.science/r/ant-main-suite-7BC0/.