🤖 AI Summary
This study addresses the security vulnerability wherein LLM agents’ persistent memory readily integrates low-trust evidence, creating unauthorized control paths. To investigate this, we propose the Backward Memory Attack (BMA), which leverages reverse planning via large language models to derive triggering memories from target actions, achieving precise, feedback-free attacks through preparation and execution phases. We introduce a novel gray-box paradigm requiring neither memory writing nor task modification, alongside the Path-CASR metric to rigorously distinguish memory-mediated behaviors from coincidental successes. Empirically, BMA attains a Macro Path-CASR of 18.8%, significantly outperforming baselines while retaining 78% efficacy across diverse strategies. Furthermore, our proposed provenance-binding authorization defense reduces the attack success rate to merely 2% while preserving 92.1% of legitimate task completion rates.
📝 Abstract
Persistent memory enables LLM agents to reuse prior experience, but creates a new security boundary: what an agent may remember is not what it should act on. We expose an unauthorized control path where edited low-trust evidence is consolidated into persistent memory, retrieved on a clean task, and used to drive a protected action. Crucially, the adversary neither writes memory nor alters the task. We introduce Backchain Memory Attack (BMA), a grey-box, LLM-driven inverse-planning attack that reasons backward from the target action to the memory that would trigger it, then to the evidence edit that would form it. BMA has two phases: preparation uses resettable trials to localize the first failed link and build experience; execution uses the frozen experience to rank and commit candidate edits without feedback. We introduce the Pathway-Certified Attack Success Rate (Path-CASR) to separate memory-mediated from coincidental hits: registered memory must form, be retrieved, drive the target behavior, and pass matched-intervention checks. Across four substrates and three decision backbones, BMA achieves 18.8% Macro Path-CASR, compared with 13.4% for the strongest access-matched baseline. Of BMA's behavioral hits, 60.3% pass all registered pathway and intervention checks versus 36.7% for the baseline. Frozen BMA edits retain 78.0% of their certified effect on average across four held-out consolidation policies. Representative memory-side controls leave 11.0% Path-CASR, whereas provenance-bound authorization reduces it to 2.0% while preserving 92.1% legitimate-action success.