Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of AI agents to indirect prompt injection in long-context scenarios, where dispersed fragments can be reconstructed into complete malicious instructions. We propose an adaptive long-context prompt injection method that introduces a novel attack paradigm combining fragmentation with adaptive search. Specifically, attack objectives are decomposed and embedded within retrieved content, while reconstruction cues are leveraged to induce agents to assemble and execute them. Furthermore, the OpenEvolve algorithm is incorporated to iteratively evolve both fragments and cues through hierarchical scoring and natural language feedback, thereby enhancing stealthiness. Experimental results demonstrate that our approach achieves a macro-average attack success rate of 61.4%, significantly outperforming baselines such as Trojan Hippo and overcoming the limitations of conventional single-instruction injection.
📝 Abstract
Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language models powering these agents also create new attack surfaces for indirect prompt injection. In particular, an attacker may not need to place a complete malicious instruction in retrieved content if the agent can reconstruct the objective from incomplete fragments distributed across a long context. In this work, we introduce adaptive long-context prompt injection (AdaLCPI), which combines long-context fragmentation with adaptive search. AdaLCPI splits an attack objective into incomplete fragments, embeds them in external content retrieved through the agent's tools, and uses a reconstruction cue to prompt the agent to combine them. It then iteratively refines the fragments and cue with OpenEvolve using graded scoring and natural-language execution feedback from the target agent. Empirically, AdaLCPI achieves higher attack success than strong adaptive baselines, reaching 61.4\% macro-average ASR compared with 32.8\% for Trojan Hippo-style and 30.0\% for AgentVigil. Safety evaluations should therefore test whether agents remain robust when harmful objectives must be reconstructed from incomplete fragments.
Problem

Research questions and friction points this paper is trying to address.

indirect prompt injection
agentic systems
long-context fragmentation
attack surface
safety evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Indirect Prompt Injection
Long-context Fragmentation
Adaptive Search
Agentic Systems
OpenEvolve