AgentTrap: Stateful Feedback Deception against Autonomous Penetration Testing Agents

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of traditional honeypots to counter the dynamic attack strategies employed by autonomous penetration testing agents. We propose the first closed-loop dynamic deception framework tailored for autonomous agents, integrating sentinel endpoint detection, application-layer stateful deception, and behavior-guided attack escalation techniques. This architecture enables continuous entrapment, controlled disclosure, and forensic evidence collection regarding agent behaviors. Experimental evaluations demonstrate that the proposed system reduces the attack success rate against genuine targets by 79.2% and successfully extracts attacker API keys in 18.8% of execution instances. These findings establish a novel paradigm for defending against AI agent-driven threats.
📝 Abstract
Autonomous penetration testing agents conduct multi-step attacks by continuously adapting their plans and actions to target responses. As a common defense, honeypots can be deployed to divert these agents from real assets by presenting decoy services, while also supporting attack tracing and active counterattacks. However, conventional honeypots rely primarily on static artifacts and predefined responses, leaving them unable to adapt to the evolving attack strategies of autonomous penetration testing agents. To this end, we present AgentTrap, the first closed-loop honeypot tailored for autonomous penetration testing agents. AgentTrap uses sentinel endpoints to avoid benign interference, stateful deception grounded in the protected application, and behavior-guided escalation to sustain engagement and collect agent-side behavioral evidence with controlled disclosures. We evaluate AgentTrap against eight autonomous penetration-testing agents in a deployed web application containing a real application endpoint and a separate honeypot endpoint configured under three defense strategies. Compared with no defense, AgentTrap reduces the aggregate real-target attack success rate from 95.8% to 79.2% and successfully elicits attacker API keys in 18.8% of the runs, outperforming static deception and fixed escalation. Furthermore, trace analysis shows that resistance to such counterattacks depends jointly on model-level recognition of deceptive requests and architecture-level isolation of sensitive resources.
Problem

Research questions and friction points this paper is trying to address.

autonomous penetration testing agents
honeypot
stateful deception
adaptive defense
Innovation

Methods, ideas, or system contributions that make the work stand out.

Closed-loop Honeypot
Stateful Deception
Autonomous Penetration Testing
Behavior-guided Escalation
Active Counterattack
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuelin Wang
Tianjin University
Jiongchi Yu
Jiongchi Yu
Singapore Management University
Software EngineeringSecurity
Y
Yanbang Sun
Tianjin University