A Step Towards Forgetting: Optimiser History and the Loss of Answer Mass

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the catastrophic forgetting induced by optimizer state histories, such as Adam momentum accumulation, which cause probabilistic erosion of prior knowledge during fine-tuning. The authors present the first decoupling of Adam’s historical states from answer quality by integrating loss decomposition, gradient age analysis, and local projection methods. This approach distinguishes confusion from leakage, revealing that early gradients are detrimental while recent gradients serve a protective function, thereby establishing a causal link between stored optimizer states and retention rates. Demonstrating that optimizer history actively erodes learned behaviors, this work proposes a momentum reset strategy that significantly reduces final loss on prior tasks and restores answer retention.
📝 Abstract
During fine-tuning, a language model can assign less probability to previously learned answers even when the current gradient acts to preserve that probability. With momentum, each update also carries gradients computed at earlier model states, and these stored contributions can push the model in the opposite direction. We investigate how this optimiser memory contributes to forgetting by separating old-task loss into confusion among its answers and leakage of probability outside the answer set. Across three language-model families, answer mass consistently declines while discrimination among old answers usually improves: the model becomes less likely to produce answers that it can still distinguish correctly. Decomposing Adam updates reveals opposing contributions to this loss of answer mass. Over training, accumulated history favours leakage, while the current gradient opposes it. Resolving history by age shows that the harmful contributions come mainly from older gradients of the new task, whereas recent gradients tend to protect the old answers. Changes in history's effect are dominated by its orientation relative to the old-task gradient. Interventions that reset momentum while matching the initial update norm establish that stored history affects retention, with state-dependent immediate effects and lower final old-task loss over longer Adam continuations, mainly through recovered answer mass. Finally, integration along finite updates shows that most sampled large loss increases are captured by local projections, while curvature along history amplifies some events. Together, these findings reveal how an optimiser's memory can erode learned behaviour even as its current gradient acts to preserve it.
Problem

Research questions and friction points this paper is trying to address.

Catastrophic Forgetting
Optimizer History
Momentum
Fine-tuning
Answer Mass
Innovation

Methods, ideas, or system contributions that make the work stand out.

Catastrophic Forgetting
Optimizer Momentum
Adam History
Answer Mass Leakage
Fine-tuning
🔎 Similar Papers
No similar papers found.