AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing benchmarks in capturing authentic privacy risks during the multi-step execution of large language model (LLM) agents. To overcome this, we construct a reproducible evaluation environment integrating real-world Model Context Protocol (MCP) tools with self-hosted services. We introduce the first trajectory-level privacy metrics and runtime auditing methodology grounded in actual service interactions, extending beyond conventional evaluations that focus solely on final outputs. Our experiments reveal significant yet previously overlooked privacy vulnerabilities in mainstream LLM agents. Furthermore, the results validate the critical importance of trajectory-level auditing for the trustworthy deployment of autonomous agents, demonstrating that intermediate interaction steps pose substantial privacy threats that final-output-only assessments fail to detect.
📝 Abstract
The rapid advancement of LLM agents has enabled systems to autonomously perform complex tasks through external tools, but their growing access to personal data introduces significant privacy risks. Existing benchmarks primarily evaluate LLM agent privacy through simulated trajectories and outcome-based metrics, limiting their ability to capture privacy risks arising during multi-step agent execution. In this work, we introduce AgentPrivArena, a framework for evaluating privacy risks in realistic LLM agent workflows. AgentPrivArena integrates authentic MCP tools and self-hosted services within a reproducible execution environment. We further propose trajectory-level privacy metrics that quantify unnecessary information access beyond final response leakage. Building on this framework, we introduce AgentPrivAudit, a runtime auditing approach for monitoring privacy violations during agent execution. Extensive experiments on state-of-the-art LLM agents reveal substantial privacy risks overlooked by existing evaluation paradigms, highlighting the importance of trajectory-level auditing for trustworthy agent deployment.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
privacy risks
multi-step execution
trajectory-level evaluation
personal data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Privacy Evaluation
Trajectory-level Metrics
Runtime Auditing
LLM Agents
MCP Tools
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Shouju Wang
Shouju Wang
Nanjing Medical University
NanomedicineMolecular ImagingRadiology
H
Haopeng Zhang
University of North Carolina at Charlotte