Learning What to Forget: Distributional Unlearning for LLM Representation Spaces

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of removing specific data distributions from large language models, where existing parametric assumptions struggle to accommodate high-dimensional representations. To this end, we propose Mamushi, a framework for distribution-level machine unlearning. This work introduces the first nonparametric selection rule based on Bayes-optimal logistic regression, which ranks forget samples via probabilistic classifiers and density ratio estimation. Furthermore, we establish non-asymptotic theoretical guarantees bounding the degradation from estimation error to the population-optimal rule. In toxic language detoxification and topic domain removal tasks, Mamushi significantly reduces the number of forget samples required to achieve target objectives, attaining superior forgetting-retention trade-offs compared to existing baselines.
📝 Abstract
Machine learning systems increasingly face the need to remove the influence of entire data domains, such as toxic language, harmful behavior, or topical content, rather than isolated records. Recent work formalizes this problem as \emph{distributional unlearning}: selecting a subset of a forget domain whose removal moves the training distribution away from an unwanted population while preserving proximity to the desired one. However, existing analyses often impose parametric assumptions to obtain tractable selection rules. These assumptions may be poorly suited to high-dimensional language-model representations. We introduce \textsc{Mamushi}, a framework for non-parametric distributional unlearning that ranks forget examples using a probabilistic classifier whose Bayes-optimal logit equals the forget-to-retain log-density ratio (up to an additive class-prior constant). We show that thresholding the population log-density ratio yields the optimal fixed-budget selection rule for our removal--preservation objective and establish a non-asymptotic transfer guarantee relating score-estimation and threshold-calibration errors to degradation from the population-optimal selection rule. Our empirical evaluation spans real-world datasets on toxic-language removal and topical-domain removal regimes using different representations, with \textsc{Mamushi} achieving a more favorable removal--preservation trade-off than other baselines. Our work shows that \textsc{Mamushi} can serve as an efficient selection approach for downstream machine unlearning procedures, reducing the number of forget examples required to reach a fixed forgetting target.
Problem

Research questions and friction points this paper is trying to address.

distributional unlearning
large language models
representation spaces
toxic language removal
machine unlearning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distributional Unlearning
Non-parametric Framework
Log-density Ratio
Representation Spaces
Machine Unlearning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pinaki Mohanty
Department of Computer Science, College of Science & College of Engineering, Purdue University, West Lafayette, IN, USA
H
Haoran Tang
Department of Computer Science, College of Science & College of Engineering, Purdue University, West Lafayette, IN, USA
M
Maggie Makar
Computer Science and Engineering Division, University of Michigan, Ann Arbor, MI, USA
Rajiv Khanna
Rajiv Khanna
Assistant Prof, PurdueCS
Machine LearningBig Data Algorithms