🤖 AI Summary
This study addresses the unclear relationship between model learning strategies and subsequent unlearning efficacy, which constrains the development of safe unlearning techniques for large models. Grounded in memorization-generalization spectrum theory, this work systematically investigates how learning dynamics influence unlearning behavior by explicitly decoupling and controlling the contributions of memorization and generalization. The methodology integrates grokking phenomenon analysis, modular addition experiments, and factual versus verbatim recall evaluations on large language models. Our findings reveal that generalization-dominant models exhibit more pronounced performance degradation on retained data during unlearning, establishing a monotonic relationship between the degree of generalization and unlearning difficulty. These insights provide critical theoretical foundations and practical guidance for optimizing safe unlearning in large models.
📝 Abstract
While unlearning seeks to negate undesired capabilities acquired through learning, little research has examined how the way models learn shapes their subsequent unlearning. In this paper, we investigate this connection from the perspectives of memorization and generalization, the two most representative yet competing strategies that models employ during training. We first classify memorization- and generalization-heavy models using grokking in modular addition and compare their responses to unlearning, showing that the latter suffer greater retain damage, i.e., a larger performance drop on the retain set. Furthermore, we conduct a finer-grained analysis by introducing bucketed modular addition, in which the respective contributions of the two strategies can be explicitly controlled across the memorization-generalization spectrum. In this setup, we reaffirm that the same trend persists and is nearly monotonic. We further demonstrate that this relationship also holds in LLM unlearning across verbatim and factual recall settings. Finally, we provide two practical insights for developing better unlearning methods, highlighting the importance of accounting for learning dynamics in unlearning.