🤖 AI Summary
This study addresses the critical security risks faced by AI agents due to external data, tool invocations, and software vulnerabilities, highlighting the urgent need for automated pre-deployment auditing. We propose a dual-role auditing architecture that decouples the Analyzer from the Exploiter, separating repository-level attack path discovery from runtime exploitation. By integrating multi-agent collaboration, static code analysis, and dynamic feedback mechanisms, the framework enables end-to-end security assessment. Additionally, we construct AgentXploit-Bench, a dedicated benchmark dataset for evaluation. Experimental results demonstrate that the proposed system achieves an end-to-end success rate of 59.3% and an injection attack success rate of 79.2%, significantly outperforming existing baseline methods.
📝 Abstract
AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as path traversal or command injection. We study authorized white-box pre-deployment auditing, where the auditor has access to the target repository and a controlled runtime, but successful attacks must still act through the task-defined attacker interface and be confirmed by an external verifier. We present AgentXploit, a two-role auditing system that separates repository-level attack-path discovery from runtime exploitation. The Analyzer Agent traces attacker-controlled inputs to sensitive operations and records code-supported candidate attack paths; the Exploiter Agent turns these paths into concrete attacks and revises them using runtime feedback. We also introduce AgentXploit-Bench, containing 72 reproducible vulnerabilities across 12 open-source AI-agent systems and frameworks. Across three runs, AgentXploit reaches 59.3% end-to-end success, compared with 38.4% for Codex. Under a token-budget-matched comparison, Codex reaches 46.3%. On AgentDojo, where injection points are provided, the Exploiter Agent reaches 79.2% attack success versus 52.7% for AgentVigil. These results highlight repository discovery and runtime exploitation as distinct challenges in end-to-end agent security auditing.