🤖 AI Summary
This study addresses the limitation of existing large language model unlearning evaluations, which focus solely on output layers while overlooking information residues in hidden states, leading to "pseudo-unlearning" where sensitive data remains extractable via probing. For the first time, this work theoretically distinguishes output suppression from representation erasure, revealing and quantifying leakage risks within hidden layers. To address this, we propose PARS, a method that employs generative probes to adversarially suppress extractable information in hidden representations, achieving genuine representation-level unlearning. Experiments demonstrate that PARS significantly reduces hidden-layer information leakage on benchmarks such as TOFU and effectively defends against relearning attacks, outperforming mainstream baselines. Ultimately, this work establishes a new paradigm for secure unlearning in large models.
📝 Abstract
Unlearning in large language models (LLMs) is typically evaluated at the output level, where a model appears to suppress sensitive or undesirable content. In this work, we show that such evaluations can create an illusion of forgetting: even when output-level leakage is eliminated, sensitive information can remain encoded in the model's hidden representations. We first provide a theoretical analysis establishing a fundamental separation between output suppression and representational erasure. Specifically, we show that the decoder can be made arbitrarily insensitive to sensitive directions, driving output-level leakage to zero, while the hidden representations retain the underlying information. To empirically validate this phenomenon, we train generative probe decoders on hidden states across transformer layers, enabling layer-wise measurement of information leakage. Across three widely used benchmarks, TOFU, MUSE, and WMDP, and state-of-the-art unlearning methods, we find that substantial sensitive information remains recoverable from hidden representations, even when standard output-level metrics indicate successful unlearning. To address this gap, we propose Probe-Adversarial Representation Suppression (PARS), an unlearning objective that adversarially minimizes the extractable information from hidden representations. PARS directly targets representational leakage and provides significantly stronger guarantees of erasure under adversarial probing and relearning attacks, outperforming all evaluated baselines. Our results highlight a fundamental limitation of existing unlearning paradigms and suggest that true forgetting in LLMs requires controlling not only model outputs, but also the information encoded in hidden representations. Codes are available at https://github.com/OptimAI-Lab/HiddenStateUnlearning.