Replay on Demand: An Emergent Curriculum for Balancing Adaptation and Forgetting in Continued Pretraining

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of balancing new domain adaptation against prior knowledge forgetting during continual pre-training by proposing an on-demand replay mechanism. Without requiring predefined data mixing ratios, this method dynamically allocates a shared training budget through learning potential assessment and forgetting monitoring. Furthermore, it constructs an online curriculum via a competitive generation strategy to achieve adaptive data replay and model state optimization. Experimental results demonstrate that the proposed mechanism significantly outperforms fixed-replay baselines across multi-scale models while effectively mitigating catastrophic forgetting. Ultimately, this work establishes a novel paradigm for resource-efficient utilization in continual learning.
📝 Abstract
Continued pretraining enables language models to adapt to new domains and knowledge, but often at the cost of forgetting previously acquired capabilities. Replay can mitigate this trade-off, but fixed replay mixtures allocate training independently of the model's actual retention needs. We introduce Replay on Demand (RoD), which instead derives the replay allocation from the model's learning dynamics. RoD jointly prioritizes adaptation samples by their remaining learning potential and replay samples by their observed forgetting. Their competition for a shared training budget yields an online curriculum that determines what to train on at each step. Across models, scales, and adaptation domains, RoD reaches or improves upon the adaptation-forgetting frontier of tuned fixed-replay baselines and model merging without prescribing a replay allocation in advance. Replay concentrates on sources that are more vulnerable to forgetting and dynamically increases and redistributes as forgetting emerges during training. Together, our results show that replay can be allocated online from the model's evolving state, targeting what is needed, when it is needed.
Problem

Research questions and friction points this paper is trying to address.

Continued Pretraining
Catastrophic Forgetting
Replay
Language Models
Domain Adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Replay on Demand
Continued Pretraining
Online Curriculum
Catastrophic Forgetting
Learning Dynamics