When Can First-Order Models of Fine-Tuning Bound Forgetting?

πŸ“… 2026-09-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of predicting and delineating the risk of forgetting specific factual knowledge during LoRA fine-tuning of large language models. By leveraging a first-order response model and finite-difference probing, we derive Freedman’s first-passage bound to quantify forgetting probability. We reveal that the failure of simplified bounds stems from response coefficient drift exceeding the boundary distance, and accordingly propose a complete bound theory grounded in drift magnitude. In preregistered experiments, this complete bound holds with 100% validity, successfully and precisely characterizing the forgetting risk for low-drift facts. This work provides a rigorous theoretical framework for understanding and controlling knowledge forgetting throughout the fine-tuning process.
πŸ“ Abstract
Fine-tuning a language model on new data can make it forget facts that it should keep. We ask whether measurements taken at the start of a fine-tuning run can bound, for each protected fact, the probability that the run makes the model forget it. In LoRA fine-tuning with stochastic gradient descent on models from 0.6B to 14B parameters, a first-order response model estimated by finite-difference probes predicts changes of per-fact margins with correlation 0.974-0.998. Predictions of forgetting built on this model nevertheless failed, because forgetting requires parameter changes far outside the region in which the model was validated. The probes can, however, bound the probability that a margin first falls below a boundary near zero: we derive Freedman and Azuma first-passage bounds for a linear surrogate of the margin and test on new runs whether they hold for the model. The bounds contain a term R that measures how much the response coefficients change during the run. The simplified Freedman bound, which sets R = 0, certified most facts but was violated in 14 of 112 conditions, and every fact on which it was violated had R>= a, where a is the distance of the fact's margin to the boundary. The complete Freedman bound certifies only facts with R<a, and it held in every condition. On the violated facts, the spread of the margin across test runs was a median of 14.6 times the prediction of the response model, so the failures are breakdowns of the model, and in our data they occurred only where R>= a. We found this pattern post hoc and tested it in two preregistered confirmatory studies with 43 new conditions: the complete bound held in all of them, and the simplified bound failed there on only 3 facts, each with R>= a. First-order models of fine-tuning can thus bound forgetting on the facts whose response coefficients change by less than their distance to the boundary.
Problem

Research questions and friction points this paper is trying to address.

fine-tuning
catastrophic forgetting
first-order models
language models
forgetting bounds
Innovation

Methods, ideas, or system contributions that make the work stand out.

First-order response model
LoRA fine-tuning
Forgetting bounds
Freedman first-passage bound
Finite-difference probes
πŸ”Ž Similar Papers
No similar papers found.