AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Models

📅 2026-03-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the vulnerability of pretrained vision-language models to adversarial perturbations under zero-shot settings, where existing adversarial fine-tuning methods often degrade cross-modal alignment and impair generalization. To mitigate this issue, the authors propose an Alignment-Guided Fine-Tuning (AGFT) framework that preserves semantic alignment between visual and textual representations while enhancing robustness. AGFT leverages the soft prediction distribution of the original model as an alignment guide, integrates text-guided adversarial training, and employs a temperature-scaling calibration mechanism. Experimental results demonstrate that AGFT substantially outperforms current approaches across multiple zero-shot benchmarks, achieving significantly improved adversarial robustness without compromising zero-shot generalization capability.

Technology Category

Computer Vision: Adversarial Attacks & RobustnessMachine Learning: Adversarial Learning & RobustnessMultiagent Systems: Adversarial Agents

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Pre-trained vision-language models (VLMs) exhibit strong zero-shot generalization but remain vulnerable to adversarial perturbations. Existing classification-guided adversarial fine-tuning methods often disrupt pre-trained cross-modal alignment, weakening visual-textual correspondence and degrading zero-shot performance. In this paper, we propose an Alignment-Guided Fine-Tuning (AGFT) framework that enhances zero-shot adversarial robustness while preserving the cross-modal semantic structure. Unlike label-based methods that rely on hard labels and fail to maintain the relative relationships between image and text, AGFT leverages the probabilistic predictions of the original model for text-guided adversarial training, which aligns adversarial visual features with textual embeddings via soft alignment distributions, improving zero-shot adversarial robustness. To address structural discrepancies introduced by fine-tuning, we introduce a distribution consistency calibration mechanism that adjusts the robust model output to match a temperature-scaled version of the pre-trained model predictions. Extensive experiments across multiple zero-shot benchmarks demonstrate that AGFT outperforms state-of-the-art methods while significantly improving zero-shot adversarial robustness.
Problem

Research questions and friction points this paper is trying to address.

zero-shot adversarial robustness
vision-language models
cross-modal alignment
adversarial fine-tuning
semantic correspondence
Innovation

Methods, ideas, or system contributions that make the work stand out.

alignment-guided fine-tuning
zero-shot adversarial robustness
vision-language models
soft alignment
distribution consistency calibration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yubo Cui
Yubo Cui
Northeastern University
3d computer visionobject trackingrobot
Xianchao Guan
Xianchao Guan
Harbin Institute of Technology, Shenzhen
artificial intelligence
Z
Zijun Xiong
Harbin Institute of Technology, Shenzhen, China
Z
Zheng Zhang
Harbin Institute of Technology, Shenzhen, China; Shenzhen Loop Area Institute, Shenzhen, China