Right Predictions, Misleading Explanations: On the Vulnerability of Vision-Language Model Explanations

📅 2026-05-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work reveals that explanation heatmaps of vision-language models (VLMs) can become misaligned with their predictions under adversarial conditions, undermining explanation reliability. To address this, the authors propose X-Shift, a novel gray-box attack that manipulates explanation heatmaps to highlight irrelevant image regions through patch-level perturbations—without altering the model’s original prediction. Notably, X-Shift requires no modification of model parameters and generalizes across multiple CLIP architectures and mainstream explanation methods. Experiments on ImageNet-1k, MS-COCO, and Flickr30K demonstrate that X-Shift significantly degrades explanation alignment under imperceptible perturbations, an effect unattainable by conventional adversarial attacks. These findings expose a fundamental vulnerability in current VLM explanation mechanisms.
📝 Abstract
Explanation mechanisms are increasingly used to support transparency and trust in vision-language models (VLMs), particularly in settings where model decisions require human oversight. However, the robustness of these explanations remains insufficiently understood. In this work, we investigate whether explanation heatmaps in VLMs, particularly CLIP-based models, faithfully reflect model reasoning under adversarial conditions. We show that explanation maps can be systematically manipulated while preserving the model's original prediction, revealing a disconnect between predictive behavior and explanation faithfulness. To study this vulnerability, we introduce X-Shift, a novel grey-box attack that perturbs patch-level visual representations to redirect explanation heatmaps toward semantically irrelevant regions without altering the predicted output. Unlike conventional adversarial attacks that aim to induce misclassification, X-Shift specifically targets the integrity of the explanation process itself. The attack operates without modifying model parameters and generalizes across multiple CLIP architectures and explanation methods. We evaluate the proposed approach on ImageNet-1k, MS-COCO, and Flickr30K, demonstrating consistent degradation in explanation alignment under imperceptible perturbations while maintaining prediction stability. Furthermore, standard prediction-oriented adversarial attacks fail to reproduce the same explanation-shifting behavior even under substantially larger perturbation budgets. Our findings highlight a fundamental limitation of current explanation mechanisms in VLMs and raise concerns about their use as reliable indicators of model trustworthiness in high-impact applications.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
explanation faithfulness
adversarial vulnerability
model transparency
explanation heatmap
Innovation

Methods, ideas, or system contributions that make the work stand out.

explanation faithfulness
vision-language models
adversarial attack
X-Shift
heatmap manipulation
🔎 Similar Papers
2024-07-30Conference on Empirical Methods in Natural Language ProcessingCitations: 0