🤖 AI Summary
This study addresses the challenge of balancing privacy protection with model utility and mitigating high computational costs in large language model (LLM) unlearning. We propose a training-free, null-space activation steering method that guides activation vectors via null-space constraints during inference. By eliminating the need for weight updates, this approach non-destructively removes sensitive knowledge, enabling plug-and-play, efficient machine unlearning. Experiments on the TOFU and MUSE benchmarks demonstrate that our method matches or surpasses existing baselines in unlearning quality while preserving general model utility with negligible degradation. This work establishes a new paradigm for low-cost machine unlearning in LLMs.
📝 Abstract
Large Language Models (LLMs) inevitably internalize substantial amounts of sensitive or private information during pre-training, while LLM unlearning aims to selectively erase specific knowledge to prevent privacy leakage with minimal loss of model utility. However, existing methods struggle to balance forget quality with utility, and typically incur substantial computational costs due to parameter fine-tuning. To address this, we propose Nullify, a training-free, non-destructive activation steering method for LLM unlearning. Nullify employs steering vectors during inference to redirect privacy-related activations away from their memorized answers, while satisfying a null-space constraint that leaves retained-query activations essentially unaffected to maintain utility. Evaluations on TOFU and MUSE show that Nullify matches or surpasses established baselines in forget quality while achieving near-lossless preservation of model utility. By avoiding weight updates entirely, Nullify serves as an efficient, plug-and-play inference-time intervention framework.