Coordinated Robustness Evaluation Framework for Vision-Language Models

📅 2025-06-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the insufficient robustness of vision-language models (VLMs) under simultaneous adversarial perturbations to both image and text modalities. To this end, we propose the first cross-modal collaborative adversarial attack framework. Our core innovation lies in constructing a general-purpose surrogate model based on multimodal fusion architectures, which explicitly models inter-modal dependencies for the first time. By jointly optimizing perturbations in a shared embedding space—through cross-modal gradient alignment and collaborative optimization—we generate coordinated image and text perturbations. Extensive experiments demonstrate that our method significantly outperforms unimodal and existing bimodal attacks on VQA and visual reasoning benchmarks. On state-of-the-art VLMs—including Instruct-BLIP and ViLT—it reduces accuracy by 32–47%. The framework establishes a unified, scalable benchmark for systematic robustness evaluation of VLMs, enabling rigorous assessment of cross-modal vulnerability.

Technology Category

Computer Vision: Adversarial Attacks & RobustnessMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Vision-language models, which integrate computer vision and natural language processing capabilities, have demonstrated significant advancements in tasks such as image captioning and visual question and answering. However, similar to traditional models, they are susceptible to small perturbations, posing a challenge to their robustness, particularly in deployment scenarios. Evaluating the robustness of these models requires perturbations in both the vision and language modalities to learn their inter-modal dependencies. In this work, we train a generic surrogate model that can take both image and text as input and generate joint representation which is further used to generate adversarial perturbations for both the text and image modalities. This coordinated attack strategy is evaluated on the visual question and answering and visual reasoning datasets using various state-of-the-art vision-language models. Our results indicate that the proposed strategy outperforms other multi-modal attacks and single-modality attacks from the recent literature. Our results demonstrate their effectiveness in compromising the robustness of several state-of-the-art pre-trained multi-modal models such as instruct-BLIP, ViLT and others.
Problem

Research questions and friction points this paper is trying to address.

Evaluating robustness of vision-language models against perturbations
Generating coordinated adversarial attacks across image and text modalities
Assessing multi-modal model vulnerabilities in visual reasoning tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generic surrogate model for joint representation
Coordinated adversarial perturbations in both modalities
Outperforms single and multi-modal attack strategies
🔎 Similar Papers
Ashwin Ramesh Babu
Ashwin Ramesh Babu
Senior Research Scientist @ Hewlett Packard Labs
Computer VisionSelf-Supervised LearningAdversarial AttacksReinforcement Learningrenewable
S
Sajad Mousavi
Hewlett Packard Enterprise (Hewlett Packard Labs)
V
Vineet Gundecha
Hewlett Packard Enterprise (Hewlett Packard Labs)
S
Sahand Ghorbanpour
Hewlett Packard Enterprise (Hewlett Packard Labs)
A
Avisek Naug
Hewlett Packard Enterprise (Hewlett Packard Labs)
A
Antonio Guillen
Hewlett Packard Enterprise (Hewlett Packard Labs)
R
Ricardo Luna Gutierrez
Hewlett Packard Enterprise (Hewlett Packard Labs)
S
Soumyendu Sarkar
Hewlett Packard Enterprise (Hewlett Packard Labs)