🤖 AI Summary
This paper exposes the illusory robustness of vision-language models (VLMs) under minimal perturbations by proposing a targeted semantic substitution attack. By aligning the token representations of source and target images within the post-merger space, the method achieves semantic tampering with imperceptible perturbation magnitudes in a white-box setting. Furthermore, it reveals a "semantic fusion" phenomenon wherein large language models rationalize contradictory visual signals into coherent narratives. Experimental results demonstrate that target semantics can be successfully injected into images at ε=2/255, achieving a complete substitution rate of 38% at ε=4/255. For videos, the substitution rate reaches 35.9% at merely ε=1/255. These findings profoundly expose critical security vulnerabilities inherent in VLMs.
📝 Abstract
Vision Language Models (VLMs) are widely deployed in safety-critical scenarios, and understanding to which extent they can be controlled by adversarial perturbation is a prerequisite for evaluating their trustworthiness. Existing representation-alignment attacks, which make a VLM perceive a target image, achieve limited success at $\varepsilon \leq 4/255$. Therefore, VLMs seems robust to perturbations in this range. We show that this robustness does not hold, as targeted semantic substitution succeeds within the same range. Specifically, we align each stream of the source image with its counterpart in the target image in the victim VLM's post-merger token space, operating under a white-box threat model. We evaluate under a strict success criterion, requiring the model to simultaneously name the target, confirm its presence, and deny the source. In images, target semantics appear at $\varepsilon = 2/255$ and complete replacement reaches 38\% at $\varepsilon = 4/255$. On video, complete replacement reaches 35.9\% at $\varepsilon = 1/255$. We also observe a phenomenon of \textit{semantic fusion}, where Large Language Model (LLM) rationalizes contradictory visual signals into a coherent narrative.