🤖 AI Summary
This work addresses the vulnerability of large language model (LLM) agents to prompt injection attacks and the limited generalization of existing red-teaming methods across new models. To overcome this, the authors propose PIMiner, a transferable red-teaming framework that requires no fine-tuning on new target models. PIMiner constructs a universal attack strategy library from scratch by training across diverse (dataset, target model) pairs, integrating agent-driven strategy learning, cross-model transfer, and few-shot query optimization. Remarkably, it achieves effective attacks on unseen models with only around ten queries. Experimental results demonstrate that PIMiner attains up to 86.7% attack success rate against Gemini-2.5-Pro, GPT-5.1, and Claude-Sonnet-4.5 on IPIArena and AgentDojo benchmarks, significantly enhancing the efficiency and practicality of cross-model prompt injection attacks.
📝 Abstract
Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing state-of-the-art prompt injection red-teaming methods primarily rely on reinforcement learning (RL), producing attacker models that often generalize poorly to new target LLMs. In this work, we develop PIMiner, an agentic system for prompt injection red-teaming. During training, PIMiner is trained on a sequence of (dataset, target model) pairs and builds a strategy library from scratch. At test time, the learned strategy library can be directly transferred to a previously unseen target LLM without additional training. PIMiner requires only a small number of queries to a target agent (e.g., 10) per test sample. Experimental results demonstrate that PIMiner achieves strong performance. On IPIArena, it attains a 76.2% ASR against Gemini-2.5-Pro, 61.9% ASR against GPT-5.1, and 42.9% ASR against Claude-Sonnet-4.5. On AgentDojo, it achieves an 86.7% ASR against Gemini-2.5-Pro, 53.3% ASR against GPT-5.1, and 40.0% ASR against Claude-Sonnet-4.5.