Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the training-inference discrepancy in Vision-Language Models (VLMs), where training losses misalign with inference outputs, and investigates their poorly understood internal mechanisms under adversarial perturbations. Using Qwen2.5-VL as a testbed, this work employs a two-stage PGD attack combined with Logit Lens, linear probing, and cross-layer rank tracking for mechanistic analysis. It provides the first mechanistic explanation of this gap, revealing that adversarial robustness stems primarily from language decoder priors rather than the vision encoder. Furthermore, it identifies novel phenomena including fixed target token ranks and active suppression of interfering signals by the decoder. Demonstrating that pixel statistics lack predictive power while linear probes achieve a separation AUC of 0.858, these findings offer new directions for deploying VLM defenses.
📝 Abstract
A targeted adversarial perturbation can drive a vision-language model's (VLM's) teacher-forced training loss for a fixed target caption to near zero, yet the same model, allowed to generate freely, produces the original, correct description with no trace of the target. We call this dissociation the train/inference gap, and give it a precise mechanistic account on Qwen2.5-VL-7B-Instruct using a controlled two-stage PGD attack on 200 held-out COCO images. First, we show that image-level pixel statistics, including a correctly re-implemented, texture-based attackability measure from the CNN robustness literature, have essentially no predictive power over which images are corrupted (best predictor r=-0.050, p=0.484; ridge regression R^2=0.069). Second, using the logit lens, we localise the gap to a single autoregressive step: the rank of the target token, conditioned on the correct first token already being generated, is fixed at exactly 3,488 out of 152,064 vocabulary entries for every image and every condition, with zero variance. Third, tracking target-token rank across all 28 LLM decoder layers reveals that the visual encoder corrupts every image's representation by a comparable margin regardless of eventual outcome, but the language model decoder then differentially arbitrates: amplifying the corrupted signal for susceptible images and actively suppressing it, past its clean-image baseline, for resistant ones (p<0.001, rank-biserial r=0.579). A linear probe on the merger hidden state separates these two outcomes with AUC=0.858, though we flag a circularity concern in this estimate. Together these results argue that adversarial robustness in autoregressive VLMs is substantially a property of the language decoder's prior, not the visual encoder, with direct implications for where faithfulness evaluations and defenses for deployed VLM systems should be targeted.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
adversarial robustness
train/inference gap
autoregressive generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

train/inference gap
adversarial robustness
vision-language models
logit lens
language decoder prior
💼 Related Jobs
No related jobs found.
A
Arun Josephraj Arokiaraj
1Department of Computer Science, University College London 2Holistic AI
Zekun Wu
Zekun Wu
Research Scientist, Holistic AI / PhD Student, University College London
Agentic AIResponsible AIBehavioural RobustnessExplainabilityInterpretability
A
Adriano Koshiyama
1Department of Computer Science, University College London 2Holistic AI