🤖 AI Summary
This study addresses the overlooked challenge of individual identity unlearning under cross-modal entanglement in vision-language models, for which existing benchmarks remain inadequate. To this end, it proposes IDUnlearn-Bench, the first instance-level multimodal unlearning benchmark. Methodologically, the work formally defines instance-level unlearning objectives to elucidate the fundamental distinction between attribute and identity unlearning, constructs interconnected multimodal evidence representations, and designs four evaluation tasks encompassing attribute retention, identity erasure, and reconstruction capability. Empirical findings reveal that attribute unlearning frequently retains residual identity information and that risk mitigation exhibits non-monotonic dynamics. These insights provide a critical evaluation foundation for safe multimodal unlearning.
📝 Abstract
Erasing individual identities from Vision-Language Models (VLMs) is uniquely challenging because personal data is entangled across modalities rather than stored as isolated attributes. However, existing multimodal unlearning benchmarks primarily evaluate attribute-centric forgetting, overlooking the more critical objective of individual-level unlearning: eliminating a model's ability to access, link, and reconstruct target-related information across modalities. To address this gap, we propose IDUnlearn-Bench, the first benchmark for individual-level multimodal unlearning in VLMs. It represents each individual as connected multimodal evidence and evaluates four task families: attribute access, identity access, identity binding, and identity reconstruction. Experiments on representative VLMs and unlearning methods show that successful attribute-centric forgetting often leaves substantial identity-level knowledge intact and can be non-monotonic: reducing one form of risk may amplify another. Models may suppress selected information while still identifying the target, linking records, or reconstructing the individual. These findings reveal a fundamental gap between forgetting information about a person and forgetting the person as a whole.