Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of open-source large language models to malicious recovery of prohibited capabilities through fine-tuning, a threat that existing retrospective unlearning methods struggle to mitigate against unknown attacks. To this end, this work proposes "preemptive unlearning," a pre-release defense mechanism that leverages gradient sealing to block internal knowledge pathways, combined with a negative-region pushing technique for ReLU activation functions, thereby preventing the illicit acquisition of knowledge via fine-tuning at its source. The primary contribution lies in establishing the first pre-release defense paradigm targeting unknown adversarial data. Extensive evaluations across multiple large language models demonstrate that the proposed approach significantly outperforms existing baselines in resisting downstream malicious exploitation, effectively ensuring model safety prior to public release.
📝 Abstract
Open-weight LLMs are released not only as fixed products but also as substrates for downstream fine-tuning. This openness, however, creates legal and ethical risks because users may misuse fine-tuning to instill illicit knowledge or enable hostile operations. Model providers therefore need apre-release defense against such acquisition, motivating the problem of preemptive unlearning. Unlike retrospective unlearning, which removes capabilities already present in a fixed model, preemptive unlearning seeks to prevent their acquisition under unseen attack data and future fine-tuning procedures. Despite its practical importance, this setting remains largely unexplored, presents distinct challenges, and is therefore the central focus of our work. We first verify that existing retrospective methods provide insufficient pre-release protection. Even when forbidden capabilities are suppressed in current outputs, forbidden-domain data can still induce gradients through internal pathways, enabling later acquisition. Motivated by this finding, we propose a gradient-sealing principle that blocks these pathways by pushing relevant pre-activations into the negative region, where ReLU-family activations exhibit zero or near-zero derivatives. Experiments across multiple LLM families demonstrate our stronger resistance to downstream acquisition than retrospective baselines, validating gradient sealing as an effective mechanism for pre-release protection.
Problem

Research questions and friction points this paper is trying to address.

Preemptive Unlearning
Open-weight LLMs
Forbidden Capability
Fine-tuning Misuse
Machine Unlearning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Preemptive Unlearning
Gradient Sealing
Open-weight LLMs
Fine-tuning Defense
Activation Manipulation
🔎 Similar Papers
No similar papers found.