PRUNE: A Patching Based Repair Framework for Certiffable Unlearning of Neural Networks

📅 2025-05-10
📈 Citations: 0
Influential: 0
📄 PDF

career value

189K/year
🤖 AI Summary
To address the high retraining cost and verification difficulty in machine unlearning, this paper proposes a lightweight, verifiable neural network patching framework for forgetting. Methodologically, it pioneers a provably correct forgetting theory grounded in the “repair” paradigm; introduces a minimal-patch strategy coupled with a gradient-guided iterative selection of representative samples to ensure strict forgetting guarantees while enhancing scalability; and integrates neural network repair, SMT-based satisfiability solving, and efficient patch optimization—enabling both single-point and batch data deletion without retraining. Experimentally, on multi-class benchmarks, the approach achieves high-fidelity forgetting (accuracy drop <1.2%), reduces verification overhead by 3–5×, and significantly lowers memory consumption compared to retraining and fine-tuning baselines.

Technology Category

Application Category

📝 Abstract
It is often desirable to remove (a.k.a. unlearn) a speciffc part of the training data from a trained neural network model. A typical application scenario is to protect the data holder's right to be forgotten, which has been promoted by many recent regulation rules. Existing unlearning methods involve training alternative models with remaining data, which may be costly and challenging to verify from the data holder or a thirdparty auditor's perspective. In this work, we provide a new angle and propose a novel unlearning approach by imposing carefully crafted"patch"on the original neural network to achieve targeted"forgetting"of the requested data to delete. Speciffcally, inspired by the research line of neural network repair, we propose to strategically seek a lightweight minimum"patch"for unlearning a given data point with certiffable guarantee. Furthermore, to unlearn a considerable amount of data points (or an entire class), we propose to iteratively select a small subset of representative data points to unlearn, which achieves the effect of unlearning the whole set. Extensive experiments on multiple categorical datasets demonstrates our approach's effectiveness, achieving measurable unlearning while preserving the model's performance and being competitive in efffciency and memory consumption compared to various baseline methods.
Problem

Research questions and friction points this paper is trying to address.

Remove specific training data from neural networks
Achieve certifiable unlearning with lightweight patches
Unlearn large data sets efficiently via representative subsets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses patching to unlearn specific data points
Certifiably guarantees targeted forgetting with patches
Iteratively unlearns representative subsets for efficiency
X
Xuran Li
Zhejiang University, Hangzhou, Zhejiang 310007, China
J
Jingyi Wang
Zhejiang University, Hangzhou, Zhejiang 310007, China
X
Xiaohan Yuan
Zhejiang University, Hangzhou, Zhejiang 310007, China
P
Peixin Zhang
Singapore Management University, Singapore 188065
Zhan Qin
Zhan Qin
Researcher, Zhejiang University
Data Security and PrivacyAI Security
Z
Zhibo Wang
Zhejiang University, Hangzhou, Zhejiang 310007, China
Kui Ren
Kui Ren
Professor and Dean of Computer Science, Zhejiang University, ACM/IEEE Fellow
Data Security & PrivacyAI SecurityIoT & Vehicular Security