๐ค AI Summary
This study addresses the catastrophic forgetting and stability-plasticity trade-off induced by non-independent and identically distributed (non-IID) data in continual learning by proposing a task vector scaling method. When encountering temporally clustered non-IID batches, this approach mitigates performance degradation through partial integration of parameter displacements while synchronously adjusting optimizer states, thereby achieving stable updates when combined with continual pre-training and post-training techniques. The authors demonstrate that intermediate scaling coefficients outperform full application, revealing that these gains cannot be replicated solely through learning rate adjustments. Experimental results indicate that the proposed method significantly enhances average model performance across diverse training scenarios, exhibiting particularly strong efficacy during prolonged sequence exposure.
๐ Abstract
A central challenge in continual learning is to acquire new knowledge without forgetting what the model has already learned. This challenge appears in language model training when training data comes from various domain-, user-, or task-specific distributions that are encountered unevenly over time. In such settings, successive minibatches are temporally clustered by distribution instead of being sampled i.i.d. from the overall data mixture. Training on temporally clustered data induces a stability-plasticity tradeoff. Adapting the model to the active distribution can improve the model on the active distribution but may lead to a performance degradation on data it previously trained on. We find that this tradeoff intensifies with longer exposure to the same distribution. We therefore ask if the parameter displacement produced by such a sequence (the task vector) should be fully retained or applied partially. We compare applying the full displacement ($\lambda=1$) with partial integration, which scales the task vector by $\lambda$ before applying it to the continuing model and scales the optimizer state by the same coefficient. Across continual pretraining, pretraining from random initialization, supervised post-training, and reinforcement post-training, we find that intermediate values of $\lambda$ often improve average continuing-model performance relative to full integration, particularly after longer same-distribution sequences. In continual-pretraining experiments with both controlled streams and naturally defined adaptation sequences, task-vector scaling outperforms full integration at the matched learning rate, showing that its benefits are not reproduced by learning-rate scaling alone.