Detecting Adversarial Images through Response Profiles of Vision-Language Models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of frozen vision-language models to adversarial image perturbations and the inadequacy of existing detection mechanisms. It proposes the first detection paradigm based on deviations in multi-prompt response distributions and their stability. Using CLIP as the backbone, the method generates responses through diverse semantic prompts, extracts statistical features to construct compact profiles, and trains a lightweight classifier to identify anomalous inputs. The proposed approach achieves strong discriminative capability across multiple adversarial attacks, significantly outperforming embedding-space geometric baselines while demonstrating superior cross-attack generalization. Ultimately, this work provides an efficient and reliable defense strategy for the secure deployment of frozen vision-language models.
📝 Abstract
Adversarial perturbations can alter the predictions of frozen vision-language models (VLMs) while leaving their confidence and image--text similarity patterns seemingly plausible. We investigate whether we can identify adversarial inputs based on the broader way an image interacts with a collection of general semantic prompts. Our detector summarizes these responses using category-level statistics, relationships among prompts, deviations from clean reference distributions, and stability under weak image transformations, producing a compact response profile that is classified by a lightweight model while the VLM remains fixed. We evaluate the approach on multiple public image datasets, several CLIP-style visual backbones, and a range of gradient-based, optimization-based, automated, and spatial attacks. The detector achieves strong discrimination in attack-specific settings and retains substantial performance when evaluated on attacks not seen during training. Under a controlled detector-specific protocol, the response-profile representation outperforms the evaluated embedding-geometry baselines. Additional analyses show that the feature groups provide complementary information and that the method remains effective under variations in the prompt configuration. We also examine inference cost and performance against detector-aware adaptive attacks. Overall, the results indicate that response patterns across semantic prompts provide a useful complementary signal for adversarial image detection in frozen VLMs.
Problem

Research questions and friction points this paper is trying to address.

adversarial detection
vision-language models
adversarial perturbations
frozen VLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial Detection
Vision-Language Models
Response Profiles
Prompt Engineering
Frozen VLM
💼 Related Jobs
No related jobs found.
A
Arash Vashagh
Trustworthy and Secure AI (TSAI) Lab, Faculty of Computer Science, University of New Brunswick, Canada
Roozbeh Razavi-Far
Roozbeh Razavi-Far
Associate Professor, University of New Brunswick; SMIEEE
Machine LearningAdversarial Machine LearningTrustworthy AIBig Data AnalyticsData Mining