🤖 AI Summary
This study addresses the vulnerability of CLIP models to embedding-space backdoor attacks, noting that existing defenses typically require white-box access or struggle to localize minute triggers. To overcome these limitations, this work proposes a lightweight black-box defense framework that precisely identifies malicious regions through segmented embedding perturbation measurement and selectively purifies only suspicious segments via semantic repair. By integrating contrastive learning with perturbation analysis, the method achieves region-level selective purification without requiring model parameters or clean reference data. Experimental evaluations demonstrate that the proposed approach reduces attack success rates to 1.05% across diverse attack scenarios while preserving 86.34% benign accuracy, significantly outperforming existing black-box defense methods.
📝 Abstract
Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities. However, recent studies reveal a critical vulnerability: embedding-space backdoor attacks. By poisoning only a tiny fraction of image--text pairs, adversaries can implant stealthy triggers that induce targeted shifts in CLIP's joint embedding space. Unlike conventional backdoors that manipulate classifier logits, these attacks corrupt representations directly, making them highly effective under extremely low poisoning ratios and difficult to detect. Existing defenses require access to model parameters, gradients, logits, or clean validation data---assumptions that rarely hold in realistic black-box deployments. Moreover, current black-box methods struggle to accurately localize small or out-of-distribution triggers. We propose CLIPGuard, a lightweight and fully black-box defense specifically designed to mitigate embedding-space backdoors in CLIP encoders. CLIPGuard identifies malicious regions by measuring segment-wise embedding perturbations and selectively purifies only suspicious segments via semantic inpainting, preserving benign visual content and alignment quality. Extensive experiments on STL-10, ImageNet, and diverse trigger families---including BadCLIP, BadNets, blended, patch-based, and typographic attacks---demonstrate that CLIPGuard reduces attack success rates to as low as 1.05% while maintaining clean accuracy up to 86.34%, consistently outperforming existing black-box defenses, including CleanCLIP and CleanerCLIP. Our code is available https://github.com/wsu-cyber-security-lab-ai/CLIPGuard.git