Selective Channel Restoration for Backdoored Vision-Language Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of vision-language model (VLM) fine-tuning to backdoor attacks and the prohibitive overhead of existing defenses by proposing PSR, a post-training defense method. This work is the first to reveal the projection vulnerability phenomenon, precisely identifying sensitive channels through sensitivity analysis of the projection layer. It further employs a perturbation-selective recovery mechanism to revert compromised parameters to their pre-trained states, enabling sparse updates. The core innovation lies in efficiently eliminating backdoors without introducing additional computational costs during inference. Experimental results demonstrate that PSR reduces attack success rates to near zero while fully preserving clean-task performance, establishing a lightweight and effective defense paradigm for the secure fine-tuning of VLMs.
📝 Abstract
Vision-language models (VLMs) exhibit strong multimodal capabilities but remain vulnerable to backdoors implanted through poisoned fine-tuning data. Existing defenses often require extensive parameter updates during fine-tuning or incur per-query overhead during inference. To address these limitations, we propose Perturb-Select-Restore (PSR), a post-training defense that performs sparse updates to the projection interface and introduces no additional computation during inference. We reveal that backdoored VLM projectors are substantially more sensitive to bounded perturbations than clean VLM projectors, a phenomenon we term projection fragility. Building on this finding, PSR identifies the output channels most sensitive to perturbations in each projection layer of a backdoored VLM and restores their parameters to the corresponding pretrained values. Experiments across multiple tasks show that PSR reduces attack success rates to near zero while preserving clean-task performance.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Backdoor Defense
Poisoned Fine-tuning
Projection Fragility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Backdoor Defense
Projection Fragility
Selective Channel Restoration
Sparse Updates
🔎 Similar Papers
2024-07-02Conference on Empirical Methods in Natural Language ProcessingCitations: 1