🤖 AI Summary
This study addresses the training-inference discrepancy in Vision-Language Models (VLMs), where training losses misalign with inference outputs, and investigates their poorly understood internal mechanisms under adversarial perturbations. Using Qwen2.5-VL as a testbed, this work employs a two-stage PGD attack combined with Logit Lens, linear probing, and cross-layer rank tracking for mechanistic analysis. It provides the first mechanistic explanation of this gap, revealing that adversarial robustness stems primarily from language decoder priors rather than the vision encoder. Furthermore, it identifies novel phenomena including fixed target token ranks and active suppression of interfering signals by the decoder. Demonstrating that pixel statistics lack predictive power while linear probes achieve a separation AUC of 0.858, these findings offer new directions for deploying VLM defenses.
📝 Abstract
A targeted adversarial perturbation can drive a vision-language model's (VLM's) teacher-forced training loss for a fixed target caption to near zero, yet the same model, allowed to generate freely, produces the original, correct description with no trace of the target. We call this dissociation the train/inference gap, and give it a precise mechanistic account on Qwen2.5-VL-7B-Instruct using a controlled two-stage PGD attack on 200 held-out COCO images. First, we show that image-level pixel statistics, including a correctly re-implemented, texture-based attackability measure from the CNN robustness literature, have essentially no predictive power over which images are corrupted (best predictor r=-0.050, p=0.484; ridge regression R^2=0.069). Second, using the logit lens, we localise the gap to a single autoregressive step: the rank of the target token, conditioned on the correct first token already being generated, is fixed at exactly 3,488 out of 152,064 vocabulary entries for every image and every condition, with zero variance. Third, tracking target-token rank across all 28 LLM decoder layers reveals that the visual encoder corrupts every image's representation by a comparable margin regardless of eventual outcome, but the language model decoder then differentially arbitrates: amplifying the corrupted signal for susceptible images and actively suppressing it, past its clean-image baseline, for resistant ones (p<0.001, rank-biserial r=0.579). A linear probe on the merger hidden state separates these two outcomes with AUC=0.858, though we flag a circularity concern in this estimate. Together these results argue that adversarial robustness in autoregressive VLMs is substantially a property of the language decoder's prior, not the visual encoder, with direct implications for where faithfulness evaluations and defenses for deployed VLM systems should be targeted.