OPC: One-Point-Contraction Unlearning Toward Deep Feature Forgetting

📅 2025-07-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing machine unlearning methods primarily adjust only the output layer, leaving sensitive information about forgotten data implicitly embedded in internal model representations—rendering them vulnerable to gradient inversion or performance recovery attacks. To address this, we propose a theoretical standard for *deep unlearning*: complete elimination of target data’s recoverability from intermediate representations via *pointwise contraction* in feature space. Based on this principle, we design the OPC (Optimal Pointwise Contraction) algorithm—the first unlearning method with provable guarantees of thorough forgetting, requiring no retraining and robust against training-free reconstruction attacks. Evaluated on multiple image classification unlearning benchmarks, OPC achieves high unlearning accuracy while significantly enhancing robustness against diverse reconstruction and recovery attacks. This work marks a paradigm shift from *superficial unlearning*—focused on output-level correction—to *essential unlearning*, ensuring irrecoverability at the representation level.

Technology Category

Machine Learning: PrivacySearch and Optimization: Learning to SearchComputer Vision: Adversarial Attacks & Robustness

Application Category

Security and Privacy: Data transparency and provenanceSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systems
📝 Abstract
Machine unlearning seeks to remove the influence of particular data or class from trained models to meet privacy, legal, or ethical requirements. Existing unlearning methods tend to forget shallowly: phenomenon of an unlearned model pretend to forget by adjusting only the model response, while its internal representations retain information sufficiently to restore the forgotten data or behavior. We empirically confirm the widespread shallowness by reverting the forgetting effect of various unlearning methods via training-free performance recovery attack and gradient-inversion-based data reconstruction attack. To address this vulnerability fundamentally, we define a theoretical criterion of ``deep forgetting'' based on one-point-contraction of feature representations of data to forget. We also propose an efficient approximation algorithm, and use it to construct a novel general-purpose unlearning algorithm: One-Point-Contraction (OPC). Empirical evaluations on image classification unlearning benchmarks show that OPC achieves not only effective unlearning performance but also superior resilience against both performance recovery attack and gradient-inversion attack. The distinctive unlearning performance of OPC arises from the deep feature forgetting enforced by its theoretical foundation, and recaps the need for improved robustness of machine unlearning methods.
Problem

Research questions and friction points this paper is trying to address.

Address shallow forgetting in machine unlearning models
Define deep forgetting via one-point-contraction criterion
Enhance resilience against performance and data reconstruction attacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

One-Point-Contraction for deep feature forgetting
Training-free performance recovery attack resilience
Gradient-inversion attack resistant unlearning algorithm
🔎 Similar Papers
2024-05-21Neural Information Processing SystemsCitations: 11
Jaeheun Jung
Jaeheun Jung
PhD candidate in Korea University
Machine learning
B
Bosung Jung
Department of Mathematics, Korea University, Seoul, Republic of Korea
S
Suhyun Bae
Department of Mathematics, Korea University, Seoul, Republic of Korea
Donghun Lee
Donghun Lee
Seoul National University, ETRI
Machine Learning