ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of inadvertently degrading benign representations during safety alignment of multimodal models by proposing the first targeted alignment framework based on modality-independent safety states. Methodologically, it introduces a novel conditional alignment mechanism governed by modality states rather than sample origins, precisely rectifying unsafe content through four-way conditional objectives. Furthermore, the authors construct ViSUv2, a large-scale dataset with independent annotations, and integrate multimodal contrastive learning with a selective safety alignment algorithm. Experimental results demonstrate that the proposed approach significantly reduces harmful outputs in cross-modal retrieval and generation tasks while fully preserving the utility of the original embedding space.
πŸ“ Abstract
Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment must suppress without unnecessarily changing benign representations. Because ethical and practical constraints prevent collecting real unsafe content at scale, existing datasets pair safe real samples with generated counterparts, but label every generated sample unsafe, even when one modality is individually safe. To address this, we introduce ShieldCLIP, the first framework to condition safety alignment on the observed safety state of each modality rather than the origin of a sample, preserving safe content while redirecting only what is unsafe. We also introduce ViSUv2, a 195k-quadruplet dataset with independent per-modality safety labels across 578 concepts and 28 categories. Using these labels, ShieldCLIP defines a four-way conditional objective beyond pair-level supervision: safe content is anchored, unsafe modalities are redirected to their safe counterparts, mixed pairs update only the unsafe branch, and coherence is enforced when both are unsafe. We evaluate ShieldCLIP on cross-modal retrieval, text-to-image generation with Stable Diffusion v1.4 and SDXL, and image-to-text generation with LLaVA. Across these settings, ShieldCLIP consistently reduces harmful outputs over prior safety-aligned encoders and strong mitigation baselines, while preserving the utility of the original embedding space. Extensive ablation studies further show that both modality-specific supervision and the selective alignment objective contribute to these gains. Source code, trained models, and ViSUv2 (under a controlled-access protocol) will be made publicly available at https://aimagelab.github.io/ShieldCLIP/.
Problem

Research questions and friction points this paper is trying to address.

multimodal foundation models
safety alignment
harmful content mitigation
modality-specific safety
CLIP
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selective Safety Alignment
Multimodal Foundation Models
Conditional Objective
Per-modality Safety Labels
Harmful Content Mitigation
πŸ”Ž Similar Papers