From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of persistent memory attacks against LLM agents, which typically prioritize injection success rates while neglecting downstream harm. To bridge this gap, this work is the first to treat attack severity as an independent optimization objective, introducing Counterfactual Memory Regret (CMR) as a quantitative metric. Methodologically, it proposes the MemHarm framework, which integrates sparse semantic editing search with an offline pairwise loss feedback mechanism, thereby enabling a paradigm shift from success-oriented to harm-oriented attacks. Experimental results demonstrate that the proposed approach significantly amplifies sustained detrimental effects on agents’ subsequent behaviors while maintaining high injection success rates, achieving state-of-the-art CMR values across multiple benchmarks.
πŸ“ Abstract
As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior. Persistent-memory attacks are typically evaluated by whether they succeed, yet successful attacks can leave persistent states with substantially different downstream consequences. We study this severity as a distinct attack-design objective and formalize it with counterfactual memory regret (CMR), the paired increase in expected downstream loss relative to clean memory. We introduce MemHarm, which predeclares a finite class of sparse, grounded semantic edits, evaluates candidates through the normal agent memory interface using offline paired-loss feedback, and certifies resolved selections within that class. Compared with attack-success optimization, CMR-guided selection produces substantially larger downstream loss while retaining most of the success-rate gain. Across two agent benchmarks and diverse memory designs, MemHarm attains the highest CMR point estimates among the evaluated general attacks on identical support. Factor-removal interventions link this harm to the selected semantic factor, and native-agent deployments verify the write-to-fresh-process attack path.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
persistent memory attacks
attack severity
counterfactual memory regret
downstream loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Memory Regret
Persistent-memory Attacks
LLM Agents
MemHarm
Attack Severity
πŸ’Ό Related Jobs
No related jobs found.
M
Mingxi Zou
Shanghai Academy of AI for Science (SAIS); Fudan University
L
Langzhang Liang
Shanghai Academy of AI for Science (SAIS); Fudan University
Z
Zhuo Wang
Fudan University
Yiyang Zhao
Yiyang Zhao
Ingdan Labs
Internet of ThingsMobile Computing
L
Lizhen Qu
Monash University
Zenglin Xu
Zenglin Xu
Fudan University
Machine LearningTrustworthy AIFederated LearningLarge Language ModelsTime Series Analysis