Visual Memory Attacks Can Persist Through The KV Cache

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability wherein adversarial images can persistently manipulate multimodal large language models via the KV cache even after being removed from the context. To this end, it proposes Persistent Visual Memory Injection (P-VMI), an attack framework that constructs malicious images through soft prompt optimization and a dedicated P-VMI image generation algorithm. By leveraging attention masking mechanisms and KV cache swapping analysis, the approach ensures that adversarial features remain resident in the KV cache after the source input is masked. This work provides the first demonstration that visual attacks can achieve long-term covert control across context removal boundaries. Evaluated on Qwen3-VL, P-VMI attains a 90% targeted success rate, with its efficacy generalizing significantly to dialogue lengths exceeding those encountered during optimization.
📝 Abstract
Modern language model systems operate autonomously over increasingly long contexts containing untrusted text and images. Can an adversarial input continue to steer a model even after that input is removed from its context? We show that attacks can be trained to persist through the key/value (KV) cache of subsequent tokens, allowing adversarial influence to outlive direct access to its source.We consider the Visual Memory Injection (VMI; Schlarmann and Hein, 2026) attack setting, in which an adversarial image that stays in the context plants a hidden backdoor: the model behaves normally until a chosen trigger elicits an attacker-chosen response. We first demonstrate persistence in this setting with optimized soft prompts, which remain effective after we mask the prompt from attention. We then introduce Persistent Visual Memory Injection (P-VMI), which optimizes images to preserve this adversarial behaviour after they are masked from attention. These attacks persist over conversations substantially longer than those used during optimization. On Qwen3-VL-8B-Instruct, P-VMI achieves up to approximately $90\%$ target success in its strongest configuration and remains effective under a stricter removal setting that exposes the image only on the first turn. A cache-swap ablation localizes the persistent influence to the KV cache. Finally, we show that these attacks can be trained to survive compaction that retains the KV cache of a summary generated by the same model, demonstrating that adversarial behaviour can persist in cached state without continued access to its source.
Problem

Research questions and friction points this paper is trying to address.

Visual Memory Attack
KV Cache Persistence
Adversarial Backdoor
Multimodal Language Models
Persistent Adversarial Influence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Persistent Visual Memory Injection
KV Cache
Adversarial Attack
Backdoor
Context Compaction
🔎 Similar Papers
No similar papers found.