Don't Forget! Decomposing the Training Dynamics of Memorization in Language Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear training dynamics of memorization mechanisms in language models by proposing a fine-grained memorization decomposition method. By decomposing loss trajectories and contrasting repeated versus rare sequences, combined with sequence-level gradient alignment analysis and parameter importance evaluation on the Pythia model family, this work systematically reveals the critical role of lower-layer parameters in memory formation, forgetting, and intervention. The research elucidates the intrinsic causes of forgetting, substantially improves memorization prediction accuracy, and successfully achieves precise parametric ablation of targeted memories. Overall, this study establishes a new paradigm for understanding and controlling the memorization behavior of large language models.
📝 Abstract
Memorization has been proposed as a mechanism to explain how language models fit the tail of their training distributions, but its training dynamics are not understood well. In this work, we take a fine-grained look at memorization by decomposing the loss trajectory of memorized sequences over training and model parameters. Across the Pythia family, we study memorization of duplicated training sequences (recitation) and rare ones (recollection). We find that memorization in both cases is characterized by sequence-level gradient alignment, though recitation suffers from misalignment with other training influences which causes forgetting, explaining the necessity for higher duplication of these examples. We further show that the lower model layers are the most involved in memorization and forgetting. Predicting memorization, our decomposition improves over a cross-entropy baseline, especially in larger models and early in training. Intervening on a small set of highly influential parameters we are able to ablate memorization in the final model. Together, these findings advance our understanding of how memorization develops during training and offer insights for predicting and intervening on it.
Problem

Research questions and friction points this paper is trying to address.

memorization
training dynamics
language models
forgetting
loss trajectory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Memorization dynamics
Loss trajectory decomposition
Gradient alignment
Forgetting
Parameter intervention