Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of balancing privacy protection with model utility and mitigating high computational costs in large language model (LLM) unlearning. We propose a training-free, null-space activation steering method that guides activation vectors via null-space constraints during inference. By eliminating the need for weight updates, this approach non-destructively removes sensitive knowledge, enabling plug-and-play, efficient machine unlearning. Experiments on the TOFU and MUSE benchmarks demonstrate that our method matches or surpasses existing baselines in unlearning quality while preserving general model utility with negligible degradation. This work establishes a new paradigm for low-cost machine unlearning in LLMs.
📝 Abstract
Large Language Models (LLMs) inevitably internalize substantial amounts of sensitive or private information during pre-training, while LLM unlearning aims to selectively erase specific knowledge to prevent privacy leakage with minimal loss of model utility. However, existing methods struggle to balance forget quality with utility, and typically incur substantial computational costs due to parameter fine-tuning. To address this, we propose Nullify, a training-free, non-destructive activation steering method for LLM unlearning. Nullify employs steering vectors during inference to redirect privacy-related activations away from their memorized answers, while satisfying a null-space constraint that leaves retained-query activations essentially unaffected to maintain utility. Evaluations on TOFU and MUSE show that Nullify matches or surpasses established baselines in forget quality while achieving near-lossless preservation of model utility. By avoiding weight updates entirely, Nullify serves as an efficient, plug-and-play inference-time intervention framework.
Problem

Research questions and friction points this paper is trying to address.

LLM unlearning
privacy leakage
forget quality
model utility
computational cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Unlearning
Training-Free
Activation Steering
Null-Space Constraint
Inference-Time Intervention
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wei Zhai
College of Computer Science and Artificial Intelligence, Fudan University
X
Xiang Liu
School of Electrical and Information Engineering, Tianjin University
Qiang Huang
Qiang Huang
Senior Lecturer, School of Computer Science, University of Sunderland
speech processingaudio signal analysisnatural language processingmultimodal information processing
R
Rui Qian
College of Computer Science and Artificial Intelligence, Fudan University
Lemao Liu
Lemao Liu
Fudan University
Large Language ModelMachine TranslationNatural Language Processing
Z
Ziwei Li
King Abdullah University of Science and Technology (KAUST), Saudi Arabia
Z
Ziqi Wang
School of Software Technology, Zhejiang University
Z
Zhitao Huang
School of Transnational Law, Peking University
D
Dejing Dou
College of Computer Science and Artificial Intelligence, Fudan University