🤖 AI Summary
Large language models (LLMs) are prone to memorizing and leaking sensitive information from their training data, while existing machine unlearning methods often compromise model utility and remain vulnerable to adversarial attacks. This work proposes Whiteout, a framework that replaces holistic unlearning with a targeted overwriting strategy. Specifically, it fine-tunes the model using carefully constructed obfuscation samples to selectively overwrite target private data, coupled with an adversarial evaluation framework to verify security. Experimental results demonstrate that this approach effectively prevents sensitive information leakage with negligible impact on general model performance. It significantly outperforms existing baselines and exhibits strong robustness against diverse black-box and white-box extraction attacks.
📝 Abstract
Modern large language models (LLMs) are trained on massive, largely unfiltered datasets, including content scraped from nearly every accessible website and user inputs. As a result, LLMs often memorize and reproduce personally sensitive information (PSI) such as birth dates, phone numbers, and home addresses. This leads to significant privacy risks, particularly for high-profile individuals such as executives, politicians, and judges. Existing mitigations largely rely on machine unlearning. However, these methods often remove more information than needed, degrade model utility and safety, and are highly vulnerable to attacks.
This paper presents Whiteout, a practical tool that, upon requests by individuals, prevents LLMs from regurgitating their genuine PSIs, by overwriting them using precise and carefully designed obfuscation samples. We evaluate Whiteout on modern LLMs of varying sizes and makers, including a widely-used OpenAI model. Results show that Whiteout effectively prevents disclosure of the targeted PSIs, has negligible impact on model utility and safety, and outperforms existing alternatives. We also test Whiteout against a wide range of countermeasures, from black-box attacks like jailbreaking to white-box adaptive attacks like relearning and quantization. Finally, we conclude with a discussion on the security and ethical implications of Whiteout.