🤖 AI Summary
This study addresses the phenomenon of pseudo-forgetting during large language model fine-tuning, investigating whether prior knowledge is merely obscured or genuinely erased. We propose a novel theoretical framework demonstrating that pseudo-forgetting arises from displacement along shared weight directions combined with normalization effects, whereas true catastrophic forgetting occurs only when individual facts are altered. These dynamics are first reproduced using a minimal associative memory model and subsequently validated through synthetic data training on Transformers alongside weight update direction removal techniques. Our experiments successfully eliminate performance collapse on synthetic data and recover ostensibly forgotten facts in pretrained language models. These findings confirm that knowledge loss induced by fine-tuning is fundamentally reversible, offering new theoretical insights into the mechanisms underlying model adaptation and knowledge retention.
📝 Abstract
Knowledge that a language model appears to forget during finetuning often remains stored and can be recovered, a phenomenon called spurious forgetting. Finetuning on new facts can even produce forgetting that undoes itself: recall of the old facts collapses, recovers as training continues on new facts alone, and only then erodes for good. We seek to understand when such forgetting is not catastrophic. A minimal associative memory reproduces these dynamics with three ingredients: keys with shared structure, concentrated new values, and normalization in the network. Finetuning moves all old representations along a common direction, hiding the old facts while preserving their relative geometry; normalization withdraws this shift once the new facts are learned, whereas fact-specific changes accumulate and cause the erosion. Moreover, subtracting the common shift eliminates the collapse in a Transformer trained on synthetic data, and removing a single direction from each weight update restores old facts in a pretrained language model. Forgetting thus combines a shared, reversible loss of access with a slow erosion of individual facts, and only the second is catastrophic. Which one dominates depends on whether the new data move old memories together or apart.