🤖 AI Summary
This study addresses the absence of standardized evaluation criteria, fragmented implementations, and insufficient robustness verification in machine unlearning for multimodal large language models. To this end, we present an open-source framework integrating data, model, and evaluation modules. Methodologically, we propose a metric meta-evaluation protocol to quantify assessment reliability, support multi-benchmark reproducibility and systematic comparison via unified interfaces, and introduce adversarial attack simulations alongside membership inference detection for robustness quantification. Experimental results demonstrate that Gradient Difference (GD) and MIP-Editor achieve superior performance, while BLEU is validated as a highly reliable evaluation metric. By bridging the gap in standardized comparative pipelines, this work significantly enhances the reproducibility and methodological rigor of unlearning evaluation.
📝 Abstract
As multimodal large language models (MLLMs) become more capable and widely deployed, concerns about privacy and safety have become increasingly pressing. Machine unlearning offers one approach to addressing these concerns by removing designated information from trained models while preserving unrelated capabilities. However, fragmented implementations and evaluation protocols, incomplete robustness testing, and limited understanding of metric reliability make progress in MLLM unlearning difficult to assess systematically. We introduce Open-MMUnlearning, an open-source, extensible framework that integrates target-model preparation, multimodal data processing, unlearning, and evaluation through shared interfaces and structured configurations. The framework supports five benchmarks spanning privacy, safety, and copyright, eight MLLMs from four model families, and twelve unlearning methods. Its evaluation suite jointly assesses forgetting effectiveness, retained utility, and robustness to model interventions, adversarial inputs, and membership inference attacks. Using a common evaluation protocol, we compare ten representative unlearning methods. In this comparison, GD and MIP-Editor tie for the highest overall score: GD achieves the highest Forget Quality, while MIP-Editor preserves more Model Utility. We further introduce a metric meta-evaluation protocol that tests faithfulness using models with controlled exposure to target knowledge and robustness under quantization and relearning. Among the thirteen evaluated metrics, BLEU achieves the highest aggregate reliability score. KS-Test attains the highest faithfulness AUC but performs less well on robustness. Together, the framework and these findings support reproducible comparison of MLLM unlearning methods and systematic assessment of evaluation reliability.