Score
Designs, implements, and proves algorithms for certified machine unlearning in continual learning systems that remove or certify removal of influence using gradient-based, Hessian/second-order, or first-order procedures and that support targeted, selective, or example-level unlearning while preserving retraining-consistent behaviour. Builds and applies multi-level evaluation and auditing methods—parameter/weight-level, pointwise/example-level, representation-level, and localization-aware analyses—and derives theoretical guarantees such as upper bounds on post-unlearning excess risk, decompositions into continual-learning excess and unlearning loss, and formal forgetting–retention trade-offs for potentially non-convex models.
This work establishes the first theoretical connection between continual learning and machine unlearning, formalizing post-unlearning performance as an excess risk minimization problem decomposed into a continual learning risk term and a forgetting loss term. Building on this framework, the study systematically adapts two theoretically grounded unlearning approaches—gradient-based and Hessian-based methods—and reveals that while the former exhibits weaker forgetting capability, it incurs nearly zero storage overhead. To balance effectiveness and efficiency, the authors propose a hybrid unlearning strategy. Theoretical analysis, supported by excess risk upper bounds for non-convex models, together with empirical validation, demonstrates that the proposed method significantly reduces storage costs while effectively preserving post-unlearning performance, thereby confirming the inherent trade-off between forgetting and knowledge retention and the efficacy of the introduced approach.
Machine unlearning for non-convex models remains challenging due to the lack of certified, efficient algorithms applicable to general non-convex loss functions. Method: This paper proposes the first first-order, black-box certified unlearning algorithm for generic non-convex losses, built upon a gradient-based backtracking retraining framework: it restores an early training checkpoint and resumes optimization solely on retained data, integrating differential privacy analysis with the Polyak–Łojasiewicz (PL) inequality to rigorously characterize certified unlearning complexity. Contribution/Results: It establishes the first $(varepsilon,delta)$-certified unlearning guarantee and generalization error bound for non-convex models without strong convexity assumptions, theoretically formalizing the trade-off among privacy, utility, and computational cost. Experiments demonstrate that the algorithm achieves significantly higher unlearning accuracy and model utility than state-of-the-art baselines in realistic privacy-sensitive settings.
Certified unlearning—guaranteeing rigorous mathematical bounds on model behavior after data removal—remains theoretically intractable for deep neural networks (DNNs) under non-convex optimization, especially when models are not fully converged. Method: This work pioneers the extension of certified unlearning to non-convex DNN training, enabling dynamic, time-sensitive unlearning requests during training. We propose an efficient gradient correction mechanism based on inverse-Hessian approximation, integrated with local stability analysis for non-convex optimization and a sequential unlearning strategy. Contribution/Results: Our approach achieves strict ε-certified unlearning while substantially reducing computational overhead compared to baselines. Experiments on three real-world datasets demonstrate significant improvements in unlearning efficiency—measured by both speed and fidelity—while provably satisfying the ε-certified unlearning guarantee. This establishes a new paradigm for practical, trustworthy machine learning systems.
This work addresses key limitations of existing data unlearning methods for machine learning models: reliance on Hessian computation, restriction to convex empirical risk minimization (ERM), and poor scalability to high-dimensional, non-convex, over-parameterized settings. We propose an efficient, provably correct online unlearning algorithm. Its core innovation is an affine stochastic recursion-based mechanism for maintaining statistical vectors that implicitly encode second-order information, enabling Newton-style online updates without explicit Hessian computation. The method operates directly on non-convergent training trajectories and provides formal unlearning certification. It reduces time and memory overhead by several orders of magnitude, achieves millisecond-scale deletion latency, and—while improving test accuracy—establishes the first theoretical bounds linking deletion capacity to generalization error.
This work addresses the “right to be forgotten” by studying efficient machine unlearning under privacy compliance: removing the influence of specified data from a trained model without full retraining, such that the post-unlearning model distribution closely approximates that of a model trained de novo on the remaining data. Methodologically, it establishes, for the first time under convexity assumptions, theoretically certified guarantees for approximate unlearning; introduces a dynamic tracking technique based on the infinite Wasserstein distance, yielding provably improved computational complexity; and integrates projected-noise SGD, differential-privacy-inspired noise injection, and convex optimization analysis. Experiments demonstrate that, under both mini-batch and full-batch settings, the proposed method reduces gradient computation cost to just 2% and 10% of the state-of-the-art, respectively, while preserving equivalent model utility and formal privacy guarantees.
This work addresses the problem of efficiently and reliably removing the memory of specific data from trained models while preserving model performance and providing provable unlearning guarantees. To this end, the authors propose a second-order certified unlearning algorithm based on uniformly convex regularization and an anisotropic Gaussian mechanism. This approach uniquely integrates uniform convexity with surrogate generalization error to analyze the optimization complexity of unlearning. Theoretically, the study shows that when the samples to be forgotten are well-predicted by the model, the corresponding optimization problem becomes easier to solve. The proposed algorithm achieves global convergence for generalized linear models—such as logistic and exponential regression—and demonstrates significant improvements over existing first-order methods.
This work addresses the significant performance degradation of existing certified unlearning methods under non-i.i.d. deletion requests, which arises from distributional shifts. To tackle this challenge, we propose the first distribution-aware certified unlearning framework tailored for non-i.i.d. deletion scenarios. Our approach leverages iterative Newton updates constrained within a trust region to effectively approximate the fully retrained model, thereby achieving efficient (ε, δ)-certified unlearning. By deriving a tighter pre-run bound on the gradient residual, the method substantially mitigates the adverse effects induced by distributional shift. Extensive experiments demonstrate that our framework consistently outperforms current certified unlearning techniques across multiple evaluation metrics in settings with distributional shift.
Existing machine unlearning algorithms lack effective means to rigorously verify whether the influence of training data has been genuinely removed. This work proposes a general auditing framework based on membership inference attacks, which leverages hypothesis testing to compute a lower bound (ε) on data-dependent unlearning parameters, thereby establishing, for the first time, empirical falsifiability of unlearning claims. The framework systematically evaluates diverse unlearning techniques—including model pruning, fine-tuning on retention sets, and rollback-based deletion—on CIFAR-100 and Shakespeare datasets. Experimental results reveal that theoretically grounded methods, such as model pruning, achieve markedly small ε bounds, whereas empirical approaches yield substantially larger ε values, highlighting a significant disparity in their practical unlearning efficacy.
This study addresses the compliance challenges posed by the "right to be forgotten" by proposing the first economic auditing framework for machine unlearning. Integrating certified unlearning theory with regulatory game-theoretic modeling, the work uniquely combines hypothesis testing interpretations with nonlinear equilibrium analysis to tackle the coupled optimization of model utility and detection probability. Through a fixed-point transformation and structured analytical approach, the authors establish the existence and uniqueness of the equilibrium solution and uncover a counterintuitive insight: higher volumes of deletion requests necessitate lower auditing intensity. Furthermore, the analysis reveals that while non-public auditing offers informational advantages, it yields inferior cost-effectiveness. These findings provide rigorous theoretical support for regulatory practices, including those in jurisdictions such as China.
This work addresses two critical challenges in sequential machine unlearning: the progressive degradation of accuracy on retained data—termed knowledge erosion—and the unintended recovery of previously forgotten samples, known as unlearning reversal. To tackle these issues systematically, the authors propose SAFER, a novel framework that enforces representational stability for retained data while introducing negative logit margin regularization for forgotten data. This dual mechanism establishes a new paradigm that simultaneously ensures effective unlearning and model stability. Experimental results demonstrate that SAFER significantly mitigates both knowledge erosion and unlearning reversal across multiple unlearning rounds, maintaining high accuracy on retained data while effectively erasing information from forgotten samples.