🤖 AI Summary
This study addresses the challenge of privacy unlearning in vision-language models, where sensitive and retained information representations are inherently entangled. To this end, we propose a multimodal selective unlearning method based on attention value intervention. By treating attention values as pivotal intervention points, our approach suppresses personally identifiable information through attention value regularization. This mechanism is further integrated with frozen reference model matching, sequence-level negative cross-entropy, and retention supervision signals to jointly optimize precise unlearning and knowledge utility. Experimental results demonstrate that the proposed method achieves state-of-the-art performance in multimodal settings, significantly mitigating utility degradation while effectively balancing privacy unlearning efficacy with the preservation of general knowledge.
📝 Abstract
The ability of vision-language models (VLMs) to associate visual identities with biographical information creates a need for selective unlearning of personally identifiable information (PII) while preserving permitted knowledge about the same individual. This setting is challenging because both sensitive and retained information can share the same visual inputs and intermediate representations. We introduce SIEVE, a simple and effective framework for selective VLM unlearning. SIEVE directly regularizes attention-value representations while also controlling model outputs. SIEVE suppresses attention values for forget examples toward a constant zero, while preserving retain-example representations by matching them to a frozen reference model. These objectives are combined with sequence-level forget and retain supervision, enabling targeted forgetting without largely affecting retained knowledge. Extensive experiments show that SIEVE achieves state-of-the-art performance on unlearning with multiple model-modality settings, while maintaining competitive retained utility. Ablation studies further show that value suppression and negative cross-entropy contribute complementary forgetting signals, while reference-based value matching substantially reduces utility degradation. These results demonstrate that attention values provide an effective intervention point for selective multimodal unlearning when sensitive and retained knowledge are closely related.