It Takes Little to Rewrite Perception: Targeted Semantic Substitution in Vision-Language Models at $ε\leq 4/255$

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper exposes the illusory robustness of vision-language models (VLMs) under minimal perturbations by proposing a targeted semantic substitution attack. By aligning the token representations of source and target images within the post-merger space, the method achieves semantic tampering with imperceptible perturbation magnitudes in a white-box setting. Furthermore, it reveals a "semantic fusion" phenomenon wherein large language models rationalize contradictory visual signals into coherent narratives. Experimental results demonstrate that target semantics can be successfully injected into images at ε=2/255, achieving a complete substitution rate of 38% at ε=4/255. For videos, the substitution rate reaches 35.9% at merely ε=1/255. These findings profoundly expose critical security vulnerabilities inherent in VLMs.
📝 Abstract
Vision Language Models (VLMs) are widely deployed in safety-critical scenarios, and understanding to which extent they can be controlled by adversarial perturbation is a prerequisite for evaluating their trustworthiness. Existing representation-alignment attacks, which make a VLM perceive a target image, achieve limited success at $\varepsilon \leq 4/255$. Therefore, VLMs seems robust to perturbations in this range. We show that this robustness does not hold, as targeted semantic substitution succeeds within the same range. Specifically, we align each stream of the source image with its counterpart in the target image in the victim VLM's post-merger token space, operating under a white-box threat model. We evaluate under a strict success criterion, requiring the model to simultaneously name the target, confirm its presence, and deny the source. In images, target semantics appear at $\varepsilon = 2/255$ and complete replacement reaches 38\% at $\varepsilon = 4/255$. On video, complete replacement reaches 35.9\% at $\varepsilon = 1/255$. We also observe a phenomenon of \textit{semantic fusion}, where Large Language Model (LLM) rationalizes contradictory visual signals into a coherent narrative.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Adversarial Perturbation
Targeted Semantic Substitution
Robustness
Trustworthiness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Adversarial Perturbation
Targeted Semantic Substitution
Post-merger Token Space
Semantic Fusion
🔎 Similar Papers
No similar papers found.