Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the suboptimality of static parameter selection in large language model unlearning and its difficulty in balancing utility against side effects by proposing a dynamic intervention re-ranking mechanism. The proposed method decouples parameter localization, initial selection, and checkpoint-dependent support revision, achieving adaptive calibration and dynamic optimization of updated parameter subsets through a designed intervention score. Furthermore, it constructs a comprehensive framework by integrating techniques such as low-rank adaptation, negative preference optimization, gradient difference objectives, and calibration probes. Experimental evaluations on the Natural-TOFU and LACUNA benchmarks demonstrate that this approach significantly enhances model utility, outperforming existing static baselines across the majority of metrics.
📝 Abstract
Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds. In a controlled experiment, a storage-localization score reaches an area under the receiver operating characteristic curve (AUROC) of 0.981, yet storage identity agrees with the better intervention on only 17/36 targets, while low-rank adaptation (LoRA) wins 35/36. We introduce Intervention Score, which ranks editable groups by the predicted effect of the actual unlearning update while accounting for collateral damage, and use it to form the static intervention-value baseline (Static-IV). We then introduce selective dynamic intervention re-ranking (DIR-R), which revisits that subset only when a calibrated probe justifies the comparison. On the Natural-TOFU dataset, our method has positive descriptive margins in 19/20 comparisons between methods and objectives, although several are near zero. On the LACUNA localization-precision benchmark, our mean terminal utility is higher in all six negative preference optimization (NPO) and SimNPO comparisons: NPO margins range from +0.431 to +0.848, and SimNPO margins range from +0.503 to +0.571. The gradient-difference (GradDiff) objective reveals substantial field dependence. Relative to Static-IV, the primary four-field GradDiff evaluation has six wins, six ties, and no losses, with mean and median paired gains of +0.165 and +0.0025. The evidence supports separating localization, initial intervention selection, and checkpoint-dependent support revision.
Problem

Research questions and friction points this paper is trying to address.

LLM unlearning
parameter localization
intervention selection
dynamic optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Unlearning
Intervention Score
Dynamic Intervention Re-ranking
State-Conditioned Support Control
Low-Rank Adaptation
🔎 Similar Papers
No similar papers found.