Render Before Reading: Visual Rendering as a Prompt Injection Defense

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the vulnerability of large language models to prompt injection attacks by identifying and quantifying a "modality asymmetry," wherein multimodal models exhibit significantly lower adherence to non-textual instructions. Leveraging this observation, we propose Pictionary, a training-free defense mechanism that renders untrusted inputs as typographic images or converts them into audio prior to inference via multimodal large language models, thereby disrupting text-level malicious injections. Experimental results demonstrate that Pictionary substantially reduces attack success rates while preserving utility on benign tasks. Furthermore, it exhibits strong resilience against adaptive adversaries. Overall, this work establishes a lightweight yet robust defensive paradigm for enhancing the security of large foundation models.
πŸ“ Abstract
Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior. In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are more likely to follow adversarial instruction when they appear as text than when the same instruction is delivered through a non-textual channel (e.g., as an image). We hypothesize that this modality gap arises from text-centric instruction tuning, which teaches models to obey textual instructions while treating other modalities mainly as content to parse or describe. We then demonstrate how this gap can be turned into a training-free defense, by rendering all untrusted payloads as typographic images (or audio) before they reach the model. Across ten models and two prompt injection benchmarks (DirectInject and AgentDojo) we show that our defense Pictionary consistently reduces attack success rates even against the strongest adaptive attacks and human red teamers, while largely preserving benign utility. We further show that benign fine-tuning on image-rendered instructions erodes the modality gap, tracing it to the text-centric instruction-tuning distribution.
Problem

Research questions and friction points this paper is trying to address.

prompt injection
multimodal LLMs
adversarial attacks
modality gap
instruction tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

prompt injection defense
modality asymmetry
training-free defense
visual rendering
multimodal LLMs