Off-Policy Merging Beats On-Policy Self-Distillation for Continual Learning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of continual learning, where training on offline data often induces catastrophic forgetting while online self-distillation frequently triggers reasoning collapse. To overcome these limitations, this work proposes a "weight grafting" strategy that integrates sensitivity-direction masking with model merging techniques. Specifically, it extracts learned updates from early checkpoints and scales them for transfer to post-trained models, enabling efficient continual learning under both SFT and RL paradigms and challenging the presumed necessity of online training. The contributions demonstrate that offline merging significantly outperforms online self-distillation while substantially reducing sampling costs. In scenarios such as expert distillation, the proposed method achieves Pareto superiority over SFT and OPSD baselines across both new and previous tasks, effectively preventing reasoning collapse while preserving generalization capabilities.
📝 Abstract
A long-standing goal of AI is a model that can continually learn and improve itself. On post-trained models, supervised finetuning (SFT) on new data often causes poor generalization and catastrophic forgetting. As such, the conventional wisdom is that on-policy training is a prerequisite for continual learning. In practice, however, data containing new knowledge or capabilities are often off-policy. While methods such as on-policy self-distillation (OPSD) try to bridge this gap by converting off-policy data into on-policy signal, they have been shown to cause reasoning collapse. In this paper, we show that off-policy merging beats OPSD for continual learning. We first show that SFT learns a useful signal from new data, but naively applying its update interferes with existing capabilities. We reduce this interference with a simple recipe we term grafting, which changes where the update is learned and how it is applied: (1) learning the update on an earlier donor checkpoint, ideally even before the end of pretraining, and applying the weight update to the post-trained model; (2) scaling the weight update, equivalent to a form of model merging; and (3) optionally, masking the most sensitive update directions when the new data distribution is far from the post-trained model. Across continual learning settings including (1) distilling from expert traces, (2) self-improvement with STaR and Pedagogical RL, and (3) injecting knowledge after pretraining cutoff, grafting Pareto-dominates both SFT and OPSD in new-task and old-task performance, while avoiding expensive on-policy sampling. Therefore, our work challenges on-policy training as a necessity for continual learning on RL-trained models.
Problem

Research questions and friction points this paper is trying to address.

Continual Learning
Off-Policy Data
Catastrophic Forgetting
Supervised Fine-Tuning
On-Policy Self-Distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continual Learning
Off-Policy Merging
Grafting
Catastrophic Forgetting
Self-Distillation