Certified Machine Unlearning via Noisy Stochastic Gradient Descent

๐Ÿ“… 2024-03-25
๐Ÿ“ˆ Citations: 1
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the โ€œright to be forgottenโ€ by studying efficient machine unlearning under privacy compliance: removing the influence of specified data from a trained model without full retraining, such that the post-unlearning model distribution closely approximates that of a model trained de novo on the remaining data. Methodologically, it establishes, for the first time under convexity assumptions, theoretically certified guarantees for approximate unlearning; introduces a dynamic tracking technique based on the infinite Wasserstein distance, yielding provably improved computational complexity; and integrates projected-noise SGD, differential-privacy-inspired noise injection, and convex optimization analysis. Experiments demonstrate that, under both mini-batch and full-batch settings, the proposed method reduces gradient computation cost to just 2% and 10% of the state-of-the-art, respectively, while preserving equivalent model utility and formal privacy guarantees.

Technology Category

Machine Learning: PrivacySearch and Optimization: Learning to SearchNatural Language Processing: Learning & Optimization for NLP

Application Category

User Modeling, Personalization and Recommendation: User privacy protection in personalized systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSecurity and Privacy: Security and privacy of machine learning and AI applications
๐Ÿ“ Abstract
``The right to be forgotten'' ensured by laws for user data privacy becomes increasingly important. Machine unlearning aims to efficiently remove the effect of certain data points on the trained model parameters so that it can be approximately the same as if one retrains the model from scratch. We propose to leverage projected noisy stochastic gradient descent for unlearning and establish its first approximate unlearning guarantee under the convexity assumption. Our approach exhibits several benefits, including provable complexity saving compared to retraining, and supporting sequential and batch unlearning. Both of these benefits are closely related to our new results on the infinite Wasserstein distance tracking of the adjacent (un)learning processes. Extensive experiments show that our approach achieves a similar utility under the same privacy constraint while using $2%$ and $10%$ of the gradient computations compared with the state-of-the-art gradient-based approximate unlearning methods for mini-batch and full-batch settings, respectively.
Problem

Research questions and friction points this paper is trying to address.

Privacy Law Compliance
Data Obliviousness
Machine Learning Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Projected Noise SGD
Machine Unlearning
Privacy-preserving Efficiency
Georgia Institute of Technology | Massachusetts Institute of Technology
Eli Chien
Eli Chien
Visiting Researcher, Google
Regulatable AIMachine UnlearningDifferential PrivacyGraph machine learning
H
Haoyu Wang
Department of Electrical and Computer Engineering, Georgia Institute of Technology, Georgia, U.S.A
Z
Ziang Chen
Department of Mathematics, Massachusetts Institute of Technology, Massachusetts, U.S.A
P
Pan Li
Department of Electrical and Computer Engineering, Georgia Institute of Technology, Georgia, U.S.A