🤖 AI Summary
This study addresses the vulnerability of machine unlearning in large language models to retraining attacks and alignment fragility by proposing the FDCU framework. This method identifies shallow alignment loopholes and introduces element-wise dual masking rules alongside a minimal functional intervention principle to thoroughly disrupt target representations. Furthermore, it achieves doubly constrained subspace projection by integrating Fisher information-based universal manifold preservation with parameter anomalous activation constraints, thereby preventing models from exploiting spurious suppression layers to conceal malicious behaviors. Experimental results demonstrate that this framework attains state-of-the-art robustness against retraining attacks in both knowledge erasure and safety control tasks while preserving nearly lossless general capabilities.
📝 Abstract
Machine unlearning has emerged as a crucial mechanism for removing hazardous knowledge and enforcing safety alignment in Large Language Models (LLMs). However, recent studies reveal a persistent security risk: unlearned models remain highly vulnerable to retraining attacks, where suppressed malicious behaviors rapidly resurface after benign fine-tuning. In this work, we investigate the optimization dynamics of unlearning and identify that this vulnerability stems from shallow alignment. Rather than effectively erasing target knowledge, models often exploit a shortcut by activating previously dormant parameters to act as spurious suppressors, forming a fragile inhibitory shell over intact malicious representations. To address this issue and enforce authentic memory deletion, we propose FDCU, a novel dual-constrained subspace projection framework. FDCU restricts parameter updates through a highly scalable, element-wise dual-masking rule: it preserves general knowledge manifolds via Fisher Information and strictly prohibits the abnormal activation of spurious suppressors via the Principle of Minimal Functional Intervention (PMFI). By reliably blocking the model's ability to superficially hide knowledge, FDCU promotes the authentic dismantling of target representations. Extensive experiments across specific knowledge erasure and safe output control tasks demonstrate that FDCU achieves state-of-the-art robustness against retraining attacks while maintaining near-lossless general utility, ensuring durable safety for LLMs.