🤖 AI Summary
This study addresses catastrophic forgetting in pretrained models adapting to new data when historical samples are scarce. To this end, we propose CPLUS, a method that innovatively leverages self-probing on frozen models to synthesize inputs and record predictions, utilizing them as gradient signals rather than conventional replay exemplars to substantially expand the evidential basis for knowledge retention. Furthermore, CPLUS incorporates gradient conflict detection with a parameter update scaling algorithm to effectively suppress destructive updates. Extensive experiments across five models and four benchmarks demonstrate that CPLUS significantly reduces forgetting rates, with its advantages becoming increasingly pronounced under severe data scarcity and at larger model scales.
📝 Abstract
Adapting pretrained models to new data can cause catastrophic forgetting of previously learned behavior. When only a few past samples remain, they give continual learning methods sparse and narrow evidence about what to preserve. We show that language models can expand this evidence through self-probing, in which the frozen model generates new inputs from the retained samples and records its own predictions on them. Unlike prior work that replays such data as training examples, our method, CPLUS uses self-probe and past-sample gradients to scale down parameter updates that conflict with prior behavior. Experiments with five language models on four benchmarks show three results. First, the same probes preserve more prior behavior as gradient signals than as replay data. Second, CPLUS learns the new data while consistently reducing forgetting more than existing baselines, especially when past data are scarce, and this protection extends to benchmarks not used for training. Third, we observe that CPLUS also becomes more effective as models grow: within the Qwen3 model family, it recovers an increasing share of the forgetting caused by standard fine-tuning.