Which Image Property Carries the Jailbreak? A Controlled Dissection of Image-to-Text Jailbreaks

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear harmful factors on the image side in vision-language jailbreak attacks and the poorly understood distribution of malicious intent across modalities. Leveraging the StrongREJECT benchmark and models such as Qwen3-VL, we conduct controlled experiments to systematically evaluate the effects of image entropy, resolution, and image-text relevance. We further rectify confounding issues in tile-counting tests by employing automated judges with rigorous statistical corrections. This work is the first to attribute attack success to specific manipulations of image-text relevance rather than mere density metrics. Our findings demonstrate that removing query relevance significantly reduces attack success rates, whereas pure density-based screening proves ineffective, thereby providing a clear attribution basis for multimodal safety defenses.
📝 Abstract
Image-to-text jailbreaks place harmful intent in text, image content, or the relationship between them. We examine image-side factors across four published attack families on a 313-prompt StrongREJECT slice, using five multimodal models and an additional appendix evaluation of InternVL3.5-8B. The harmful instruction is held constant across conditions; the baseline matrix uses one draw per prompt, and paired ablations use three draws with an automated rubric judge. A bare harmful query, with or without a benign unrelated image, produces little attack success on most victims, while attack images substantially increase it. First, per-tile entropy and JPEG size do not reliably distinguish attack tiles from size-matched benign distractors, limiting density-only screening. Second, earlier tile-count ladders were confounded by payload visibility. A corrected region-count test found no detectable effect, so the role of tile-count structure remains unresolved. Third, on Qwen3-VL-8B, the E4 manipulation that removes query-specific relatedness lowers ASR by about 0.12. This supports a bounded attribution to the relatedness manipulation, although image-text congruence remains unmeasured. A within-category control reproduces the direction on overlapping stimuli. Five tests survive the global statistical correction, but only E4 supports attribution to one measured descriptor; the within-category result is a robustness check, not a separate attribution. These conclusions remain conditional on the rubric judge.
Problem

Research questions and friction points this paper is trying to address.

multimodal jailbreak
image-to-text attack
image properties
vision-language models
attack attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Jailbreak
Controlled Ablation
Image-to-Text Attack
Per-tile Entropy
Automated Rubric Judge
B
Boyuan Chen
eBRAIN Lab, Division of Engineering, New York University Abu Dhabi, UAE
Y
Yehia Dawoud
eBRAIN Lab, Division of Engineering, New York University Abu Dhabi, UAE
H
Hailemariam Mersha
eBRAIN Lab, Division of Engineering, New York University Abu Dhabi, UAE
M
Minghao Shao
eBRAIN Lab, Division of Engineering, New York University Abu Dhabi, UAE
Siddharth Garg
Siddharth Garg
Institute Associate Professor, New York University
AI/MLHardwareSecurityPrivacy
R
Ramesh Karri
Tandon School of Engineering, New York University, NY, USA
Muhammad Shafique
Muhammad Shafique
Professor, ECE, New York University (AD-UAE, Tandon-USA), Director eBRAIN Lab
Embedded Machine LearningBrain-Inspired ComputingRobust & Energy-Efficient System DesignSmart