π€ AI Summary
This study addresses the limitation of persistent memory attacks against LLM agents, which typically prioritize injection success rates while neglecting downstream harm. To bridge this gap, this work is the first to treat attack severity as an independent optimization objective, introducing Counterfactual Memory Regret (CMR) as a quantitative metric. Methodologically, it proposes the MemHarm framework, which integrates sparse semantic editing search with an offline pairwise loss feedback mechanism, thereby enabling a paradigm shift from success-oriented to harm-oriented attacks. Experimental results demonstrate that the proposed approach significantly amplifies sustained detrimental effects on agentsβ subsequent behaviors while maintaining high injection success rates, achieving state-of-the-art CMR values across multiple benchmarks.
π Abstract
As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior. Persistent-memory attacks are typically evaluated by whether they succeed, yet successful attacks can leave persistent states with substantially different downstream consequences. We study this severity as a distinct attack-design objective and formalize it with counterfactual memory regret (CMR), the paired increase in expected downstream loss relative to clean memory. We introduce MemHarm, which predeclares a finite class of sparse, grounded semantic edits, evaluates candidates through the normal agent memory interface using offline paired-loss feedback, and certifies resolved selections within that class. Compared with attack-success optimization, CMR-guided selection produces substantially larger downstream loss while retaining most of the success-rate gain. Across two agent benchmarks and diverse memory designs, MemHarm attains the highest CMR point estimates among the evaluated general attacks on identical support. Factor-removal interventions link this harm to the selected semantic factor, and native-agent deployments verify the write-to-fresh-process attack path.