Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large vision-language models are vulnerable to backdoor attacks, yet existing defenses incur prohibitive costs by relying on full retraining or inference-time intervention. This work proposes OrthoPurify, revealing that backdoors are encoded within a small number of hijacked directions in the weight updates. The method fine-tunes the model using only a few clean samples to construct a pseudo-benign reference model, and precisely eliminates the hijacked directions via single-step orthogonal projection, thereby avoiding both full retraining and inference overhead. Experimental results demonstrate that OrthoPurify reduces the attack success rate to near zero while fully preserving the model's original performance, achieving efficient and low-cost backdoor purification.
📝 Abstract
Large vision-language models (LVLMs) are increasingly deployed in safety-critical applications, yet they remain vulnerable to backdoor attacks. Defending against such attacks remains costly, as existing methods require either extensive retraining on clean data or per-query intervention at inference time. To address this limitation, we propose OrthoPurify, a more efficient method to purify backdoored model weights via one-step orthogonal projection. Specifically, through structural analysis of backdoor weight updates, we find that the backdoor is encoded by diverting a small number of weight update directions from task adaptation to backdoor shortcut encoding, a phenomenon we term direction hijacking. However, identifying these hijacked directions requires a benign reference model, which is typically inaccessible to the defender. We show that a pseudo-benign model, obtained by fine-tuning the pretrained weights on only a small set of clean samples, provides a sufficient approximation, as the dominant update directions stabilize within the first few gradient steps. OrthoPurify uses this pseudo-benign reference to isolate the hijacked directions and removes them through a single projection on the weight update. Extensive experiments show that OrthoPurify reduces the attack success rate to near zero while preserving the original performance across diverse benchmarks, without retraining the backdoored model or introducing inference-time overhead. Our code is publicly available at https://github.com/womeimingzi/OrthoPurify.
Problem

Research questions and friction points this paper is trying to address.

backdoor attacks
large vision-language models
model purification
defense efficiency
direction hijacking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Backdoor Defense
Large Vision-Language Models
Orthogonal Projection
Direction Hijacking
Weight Purification