ASPIRE: Agentic Safety & Prompt Injection Red-teaming Engine

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the risk of indirect prompt injection in LLM agents caused by retrieving untrustworthy content, as well as the limitation of existing red-teaming methods in exploring latent vulnerabilities within behavioral spaces. To this end, it proposes an open-ended, behavior-level vulnerability discovery engine. The framework constructs an agent safety behavior graph and employs a complementary dual-expert strategy balancing exploration and exploitation. By introducing a trajectory-evidence-based diagnostic mechanism and cross-run memory transfer, it transcends predefined scenario constraints to enable end-to-end attack exploration, validation, and generalization. Experimental results demonstrate that the proposed method significantly expands coverage across consequences, injection techniques, and behavioral paths on multiple benchmarks while maintaining high attack success rates.
📝 Abstract
LLM agents retrieve untrusted content and act through tools, creating indirect prompt-injection risks that can cause unauthorized actions or persistent state changes. Existing automated red-teaming largely optimizes payloads for pre-specified scenarios, leaving latent vulnerabilities across the agent's behavior space unexplored. We present ASPIRE, an Agentic Safety & Prompt Injection Red-teaming Engine for open-ended, behavior-level vulnerability discovery. ASPIRE maintains an evolving Agent Security Behavior Graph and uses complementary Explore and Exploit experts to discover, verify, and generalize consequence-centric tests. Trajectory evidence updates the graph and diagnoses partial or failed attempts, while cross-run memory transfers useful red-team strategies. Experiments on various benchmarks show that ASPIRE substantially expands coverage across consequences, injection methods, environments, and behavior paths while maintaining strong attack success.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
indirect prompt injection
automated red-teaming
agentic safety
vulnerability discovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prompt Injection Red-teaming
Agentic Safety
Behavior Graph
Cross-run Memory
LLM Agents