🤖 AI Summary
This study addresses the vulnerability of existing training-based defenses against multi-turn covert indirect prompt injection, which stems from their reliance on static attack distributions. We propose a co-evolutionary defense framework that employs dynamic adversarial search to continuously uncover vulnerabilities, transforming defense failures into a dynamic curriculum that overcomes the limitations of static training. By integrating LoRA fine-tuning, the GDPO algorithm, and a decoupled reward design encompassing safety, task utility, and formatting constraints, the approach guides models to learn deep semantic boundaries rather than superficial cues. Experimental results demonstrate that our method reduces the attack success rate by 88.5% across multiple benchmarks while improving overall performance by 38.0% compared to baselines, significantly enhancing the instruction robustness of AI agents.
📝 Abstract
Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based defenses are typically optimized on a static distribution of explicit injections. They learn surface-form cues rather than the boundary between serving the user and obeying an injected objective, and therefore fail when malicious intent is folded into a plausible workflow and deferred for several turns. We present CoDeL, a defense that hardens agent against an attack distribution it reshapes as it trains. The defender is updated each round via LoRA-based GDPO under a decoupled reward over safety, task progress, and format compliance, so refusing injections and completing the user's task jointly define fitness. To keep supplying it with the failures worth learning from, a co-evolving prober searches over injection rounds, attack methods, and payloads for injections that still penetrate the current defender, guided jointly by attack success and attack latency so that it preferentially mines breaches the defender notices too late. Each defender update invalidates part of the attack population and forces the next round onto a new frontier, turning the defender's own failures into a moving curriculum. Extensive experiments on three IPI benchmarks, nine baselines, and two base models show that CoDeL reduces attack success rate (ASR) by 88.5% and outperforms other baselines largely (+38.0%). Codes are available.