🤖 AI Summary
This study addresses the vulnerability of LLM agents to evolving threats such as jailbreaks and prompt injections, against which existing defenses struggle to adapt. We propose a training-free, self-evolving defense framework that distills harmful interaction trajectories into reusable safety policies via trajectory distillation. By leveraging retrieval-augmented generation, the framework achieves continuous adaptive protection without model retraining, effectively balancing safety and task utility. Experimental results demonstrate that our approach reduces the success rate of targeted injections to 0.42% on the AgentDojo benchmark, while limiting the success of adaptive attacks on HarmBench to merely 7.8%, significantly outperforming existing methods.
📝 Abstract
Large language models (LLMs) increasingly power agents that access sensitive information, use external tools, and modify software repositories. Although these capabilities offer substantial benefits, they also create security risks such as jailbreaks, prompt injection, and vulnerable code generation. Existing defenses often require retraining, fail to adapt to evolving attacks, or address only a single threat pattern. To address these limitations, we propose Self-Evolving Defense (SED), a training-free framework that distills harmful agent trajectories into reusable security policies without updating model weights. By retrieving relevant policies for future tasks, SED continually adapts to new attacks while retaining knowledge across attack scenarios. To evaluate the effectiveness of SED, we test it with three open-source models (DeepSeek V4 Flash, GLM 5.2, and Kimi K3) on eight benchmarks that span jailbreaks, prompt injection, and insecure code generation. SED lowers targeted prompt-injection success on AGENTDOJO to 0.42%, compared with 3.7% for the best baseline defense, and holds adaptive X-TEAMING attack success on HARMBENCH to 7.8%, more than four times lower than the best baseline at 35.2%, while preserving benign task utility.