MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval

📅 2025-12-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work identifies a novel trust vulnerability at the intersection of long-term memory in LLM-based agents and Retrieval-Augmented Generation (RAG): adversaries can inject a small number of semantically plausible yet behaviorally malicious task templates—crafted as “successful” experiences—into the agent’s experience repository. Leveraging the agent’s inherent *semantic imitation heuristic* (i.e., preference for reusing retrieved, similar successful patterns), such poisoned experiences persistently hijack reasoning across sessions. Unlike prompt injection or factual knowledge corruption, this attack introduces the first *experience-poisoning* paradigm—an indirect behavioral manipulation strategy. Evaluated on MetaGPT/GPT-4o and DataInterpreter agents, the method combines lexical-embedding hybrid retrieval with RAG storage manipulation. With only 3–5 poisoned records injected, malicious experiences dominate retrieval (exceeding 70% semantic similarity weight) for analogous tasks, inducing significant yet stealthy behavioral deviations.

Technology Category

Multiagent Systems: Adversarial AgentsMachine Learning: Large Multimodal Models (LMMs)Search and Optimization: Learning to Search

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Agentic searchUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systems
📝 Abstract
Large Language Model (LLM) agents increasingly rely on long-term memory and Retrieval-Augmented Generation (RAG) to persist experiences and refine future performance. While this experience learning capability enhances agentic autonomy, it introduces a critical, unexplored attack surface, i.e., the trust boundary between an agent's reasoning core and its own past. In this paper, we introduce MemoryGraft. It is a novel indirect injection attack that compromises agent behavior not through immediate jailbreaks, but by implanting malicious successful experiences into the agent's long-term memory. Unlike traditional prompt injections that are transient, or standard RAG poisoning that targets factual knowledge, MemoryGraft exploits the agent's semantic imitation heuristic which is the tendency to replicate patterns from retrieved successful tasks. We demonstrate that an attacker who can supply benign ingestion-level artifacts that the agent reads during execution can induce it to construct a poisoned RAG store where a small set of malicious procedure templates is persisted alongside benign experiences. When the agent later encounters semantically similar tasks, union retrieval over lexical and embedding similarity reliably surfaces these grafted memories, and the agent adopts the embedded unsafe patterns, leading to persistent behavioral drift across sessions. We validate MemoryGraft on MetaGPT's DataInterpreter agent with GPT-4o and find that a small number of poisoned records can account for a large fraction of retrieved experiences on benign workloads, turning experience-based self-improvement into a vector for stealthy and durable compromise. To facilitate reproducibility and future research, our code and evaluation data are available at https://github.com/Jacobhhy/Agent-Memory-Poisoning.
Problem

Research questions and friction points this paper is trying to address.

Persistent compromise of LLM agents via poisoned long-term memory
Exploits semantic imitation heuristic to induce unsafe behavioral patterns
Demonstrates stealthy attack through poisoned RAG store in agent workflows
Innovation

Methods, ideas, or system contributions that make the work stand out.

Poisoning agent memory via malicious experience injection
Exploiting semantic imitation heuristic for persistent compromise
Inducing behavioral drift through poisoned RAG store retrieval
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.