CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of existing training-based defenses against multi-turn covert indirect prompt injection, which stems from their reliance on static attack distributions. We propose a co-evolutionary defense framework that employs dynamic adversarial search to continuously uncover vulnerabilities, transforming defense failures into a dynamic curriculum that overcomes the limitations of static training. By integrating LoRA fine-tuning, the GDPO algorithm, and a decoupled reward design encompassing safety, task utility, and formatting constraints, the approach guides models to learn deep semantic boundaries rather than superficial cues. Experimental results demonstrate that our method reduces the attack success rate by 88.5% across multiple benchmarks while improving overall performance by 38.0% compared to baselines, significantly enhancing the instruction robustness of AI agents.
📝 Abstract
Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based defenses are typically optimized on a static distribution of explicit injections. They learn surface-form cues rather than the boundary between serving the user and obeying an injected objective, and therefore fail when malicious intent is folded into a plausible workflow and deferred for several turns. We present CoDeL, a defense that hardens agent against an attack distribution it reshapes as it trains. The defender is updated each round via LoRA-based GDPO under a decoupled reward over safety, task progress, and format compliance, so refusing injections and completing the user's task jointly define fitness. To keep supplying it with the failures worth learning from, a co-evolving prober searches over injection rounds, attack methods, and payloads for injections that still penetrate the current defender, guided jointly by attack success and attack latency so that it preferentially mines breaches the defender notices too late. Each defender update invalidates part of the attack population and forces the next round onto a new frontier, turning the defender's own failures into a moving curriculum. Extensive experiments on three IPI benchmarks, nine baselines, and two base models show that CoDeL reduces attack success rate (ASR) by 88.5% and outperforms other baselines largely (+38.0%). Codes are available.
Problem

Research questions and friction points this paper is trying to address.

Indirect Prompt Injection
LLM-based Agents
Training-based Defense
Multi-turn Attack
Innovation

Methods, ideas, or system contributions that make the work stand out.

Indirect Prompt Injection
Co-evolutionary Defense
LLM-based Agents
Decoupled Reward Optimization
Adversarial Probing
🔎 Similar Papers
2024-02-20Conference on Empirical Methods in Natural Language ProcessingCitations: 8
X
Xiao Yang
School of Computer Science and Engineering, Beihang University; State Key Laboratory of Complex & Critical Software Environment, Beihang University
Y
Yangchen Ou
School of Computer Science and Engineering, Beihang University
Y
Yuhan Gao
School of Computer Science and Engineering, Beihang University
Le Wang
Le Wang
Beihang University
AgentTrustworthy AI
Zonghao Ying
Zonghao Ying
SKLCCSE, BUAA
Trustworthy AI
A
Aishan Liu
School of Computer Science and Engineering, Beihang University; State Key Laboratory of Complex & Critical Software Environment, Beihang University