Exploring and Bridging Knowledge Holes in Unlearned Multimodal Large Language Models

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical limitation in current unlearning methods for multimodal large language models (MLLMs): the inadvertent degradation of performance on semantically related benign inputs when removing harmful content—a phenomenon termed “knowledge voids”—which existing evaluation protocols fail to adequately capture. The study systematically characterizes and quantifies this issue, introduces a dedicated benchmark for probing knowledge voids, and proposes Selective Protection with Anchored Regularization (SPAR). SPAR integrates anchored activation filtering and entity abstraction enhancement to precisely excise target content while preserving general semantic structures. Experiments demonstrate that SPAR achieves a 0.00% attack success rate while recovering over 98% of original response quality, substantially outperforming existing baselines (which retain less than 50%) without compromising overall model utility.
📝 Abstract
Machine unlearning offers a promising approach to remove unsafe content from Multimodal Large Language Models (MLLMs), yet ensuring the precision of unlearning remains a persistent challenge. One reason is that current MLLM unlearning evaluation paradigms suffer from a critical blind spot: they assess model utility through benchmarks whose representations are distant from the forget set, failing to capture knowledge holes---severe degradation on benign adjacent inputs. To probe knowledge holes in unlearned MLLMs, we construct a benchmark that captures unintended degradation on benign inputs sharing generic patterns with the forget set, and confirm through controlled experiments that they are a systematic consequence of commonly used approaches. Furthermore, to bridge this gap, we propose Selective Protection with Anchored Regularization, which protects generic patterns via anchored activation filtering while reinforcing them through entity-abstracted enhancement. Our experiments on SafeEraser demonstrate that SPAR recovers over 98% of vanilla response quality compared to below 50% for standard baselines---while achieving 0.00% attack success rate and competitive model utility. These results underscore the necessity of more fine-grained evaluation for trustworthy MLLM unlearning.
Problem

Research questions and friction points this paper is trying to address.

machine unlearning
multimodal large language models
knowledge holes
evaluation blind spot
model utility
Innovation

Methods, ideas, or system contributions that make the work stand out.

machine unlearning
multimodal large language models
knowledge holes
anchored regularization
selective protection
🔎 Similar Papers
No similar papers found.
J
Junxiang You
University of Chinese Academy of Sciences
J
Junkai Chen
Institute of Automation, Chinese Academy of Sciences
Y
Yuhao He
Institute of Automation, Chinese Academy of Sciences
Ruiqi Liu
Ruiqi Liu
Texas Tech University
nonparametric methodsmachine learningeconometrics
Z
Zhetao Guo
Cloudspace Technology
S
Shu Wu
Institute of Automation, Chinese Academy of Sciences