🤖 AI Summary
This study investigates the unclear dynamic mechanisms underlying self-distillation and supervised fine-tuning during the continuous reasoning of large language models. By leveraging directed acyclic graph (DAG) modeling and optimization-theoretic analysis, this work provides a unified characterization of post-training dynamics alongside the influence of pretraining. Specifically, the research reveals that the gradient sparsity discrepancy between online preference self-distillation (OPSD), which mitigates forgetting, and supervised fine-tuning (SFT), which induces catastrophic forgetting, constitutes the core mechanism driving these divergent behaviors. Furthermore, this project establishes theoretical guarantees demonstrating that continuous reasoning fundamentally depends on the interplay between parameter updates and pretrained structural priors. Ultimately, these findings offer critical theoretical foundations and methodological guidance for achieving efficient continuous learning in large-scale models.
📝 Abstract
On-policy self-distillation (OPSD) of large language models (LLMs) has demonstrated the ability to improve reasoning capabilities while preserving previously acquired knowledge. Despite substantial empirical success, the dynamics of OPSD in continual reasoning remain incompletely understood. Modeling LLM reasoning as search over a directed acyclic graph, we provide a unified theoretical analysis of both the dynamics of post-training---OPSD and supervised fine-tuning (SFT) in continual learning---and the impact of pre-training on subsequent performance. Our findings establish three key insights with an optimization guarantee: (i) OPSD with hints from correct outputs enables continual learning without forgetting by sparse yet effective gradient descent updates induced by the hint structure. (ii) SFT on correct reasoning paths can lead to catastrophic forgetting due to dense updates along the training paths, which overwrite the information previously acquired. (iii) Diversity in pre-training is crucial for enabling a post-trained model to reach a correct output when a rollout starts from an intermediate state. Our results, supported by theoretical analysis, show that reliable continual reasoning depends on how post-training updates interact with the reasoning structure established during pre-training.