The Uncontrolled Variable: Vision-Language Model Refusal Responds to Image Presence in Ways Risk Cannot Explain

📅 2026-08-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a vulnerability in aligned vision-language models (VLMs), where the mere presence of an image attachment—even a blank one—erroneously shifts refusal thresholds, causing benign requests to be incorrectly rejected. Through black-box testing, ablation studies, cross-checkpoint comparisons, and pixel-level perturbation analysis, this work provides the first demonstration that such refusal behavior is driven by the formal presence rather than the semantic content of the image. Furthermore, it quantifies the safety performance fluctuations and cost mismatches induced by irrelevant image attributes. Experiments reveal that blank canvases increase benign refusal rates by 23 to 51 percentage points in certain models, a bias that varies or even reverses with changes in color and resolution, and remains unresolvable through explicit instructions.
📝 Abstract
Vision-Language Model (VLM) safety is expected to depend on what a request asks for. We show that safety-aligned VLMs also key refusal on a property of a request's form: whether an image is attached, holding everything the request asks fixed. Attaching a blank canvas - unreadable, unrelated to the request, identical across prompts - shifts refusal by tens of percentage points, with no defense in the loop. The shift is not blanket caution. Neutral instructions are almost unaffected while borderline-benign prompts move sharply, so the cost falls on sensitivity-adjacent traffic: benign questions about privacy, self-harm and violence. Attachment alone is sufficient, while the image's properties set the price: a black canvas costs substantially more than a white one of identical size, and on an open checkpoint the carrying axis is pixel count. Nor is the shift under instructional control - telling the model the image is a placeholder to be disregarded removes only a fraction of it, and on one model asserting that an attachment exists moves refusal substantially with nothing attached. Attachment may correlate with risk in deployment; what these models do with it does not track risk. It is not the serving stack, since the same weights reached two ways behave alike, nor a property of VLMs as such, since several open-weight checkpoints show nothing. It belongs to particular aligned checkpoints, one of them open. It is also decoupled from what it buys: the canvas does prevent some attack success on a matched harmful set, but far less than it costs, and its sign is not fixed - on one open model the identical canvas makes the model markedly easier to attack. Image presence is not a default a deployer chose or priced; it is an uncontrolled variable inherited with the weights.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Refusal Behavior
Safety Alignment
Robustness
Over-refusal
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Safety Alignment
Refusal Behavior
Robustness
Uncontrolled Variable
💼 Related Jobs
No related jobs found.
H
Haoyu Zhang
Y
Yi Feng
H
Hanwen Liu
S
Shibo Zheng
Z
Zhuoxi Wang
Y
Yang Chen
H
Haowen Xu
X
Xiangchen Guan
M
Mohammad Zandsalimy
Shanu Sushmita
Shanu Sushmita
Northeastern University
Information Retrieval and Machine Learning