eval-unlearn: Benchmarking unlearning in Text-to-Image Diffusion Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of standardized evaluation criteria and the difficulty of cross-method comparison in concept unlearning for text-to-image models by proposing an open-source, unified benchmark framework. Built upon a plug-and-play architecture enabling automatic method registration, the framework integrates mainstream approaches including fine-tuning, closed-form editing, and inference-time intervention. It further establishes multidimensional evaluation metrics alongside a streaming batch-processing pipeline to facilitate efficient reproducibility. Accompanied by a public leaderboard and real-time interactive evaluation tools, the framework provides a fair and comparable assessment environment for diverse methods. Experimental results reveal a significant trade-off between forgetting accuracy and generation quality. Ultimately, this study offers essential infrastructure for standardizing research in machine unlearning for generative models.
📝 Abstract
The rising number of concept unlearning techniques for text-to-image (T2I) diffusion models has produced a fragmented evaluation landscape. Methods are assessed under heterogeneous experimental conditions making principled cross-method comparison difficult. We present eval-unlearn, an open-source Python library providing a unified, reproducible benchmarking framework for concept unlearning in T2I Diffusion models. eval-unlearn integrates twelve published unlearning techniques spanning fine-tuning, closed-form model editing, and inference-time intervention, alongside nine complementary evaluation metrics covering erasure efficacy, adversarial robustness, generative quality, and concept retention. Its plugin architecture lets third-party techniques and metrics self-register without modifying the core framework, and its streaming, batched pipeline supports efficient evaluation of both standard NSFW concepts and arbitrary general concepts. As a further contribution, we release a public leaderboard on HuggingFace along with an interactive tool for real-time evaluation of unlearning techniques. The leaderboard compares nudity concept erasure case study across all twelve techniques, exposing significant accuracy-quality trade-offs that are obscured by heterogeneous evaluation. eval-unlearn is released under the MIT license; the package, code, leaderboard, and documentation are all available at https://eval-unlearn.readthedocs.io.
Problem

Research questions and friction points this paper is trying to address.

concept unlearning
text-to-image diffusion models
benchmarking
evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Concept Unlearning
Text-to-Image Diffusion Models
Benchmarking Framework
Plugin Architecture
Evaluation Metrics
M
Mansi
Imperial College London
N
Nikhil Raghavan
Imperial College London
Z
Zixia Huang
Imperial College London
K
Kai Sheng Ong
Imperial College London
J
Ji Shen Lim
Imperial College London
B
Brandon Siao Xiang Ling
Imperial College London
Francesco Leofante
Francesco Leofante
Imperial College London
Artificial Intelligence