🤖 AI Summary
This study addresses the limited generalization capability of Supervised Fine-Tuning (SFT) and its neglect of parameter update directions by proposing OPSFT. The core innovation lies in, for the first time, reformulating on-policy cumulative update directions as explicit optimization constraints for SFT. This reveals how dynamically adjusting update directions facilitates generalization, thereby enabling cross-paradigm transfer of generalization capabilities. By integrating the computational efficiency of SFT with the high-quality trajectory advantages of on-policy methods, OPSFT unifies both paradigms. Experimental results demonstrate that OPSFT significantly enhances both the generalization performance and training efficiency of SFT. Furthermore, it supports the continuous improvement of large language model capabilities through the incorporation of high-quality data.
📝 Abstract
The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for improving the generalization of other paradigms such as supervised fine-tuning (SFT). To address this limitation, we investigate whether there exists a specific on-policy update behavior that can achieve such improvements. First, our analyses reveal that SFT updates parameters along consistent directions, while the on-policy paradigm continuously adjusts the direction during training. This difference inspires us to focus on the cumulative update direction of each parameter as a promising behavior. Then, we evaluate its effectiveness for improving generalization by proposing On-Policy direction-constrained Supervised Fine-Tuning (OPSFT), which constrains SFT updates to the direction identified by on-policy paradigms. The strong performance of OPSFT indicates that the generalization advantage of on-policy paradigms can be transferred to SFT through the parameter update direction. Once such a direction is identified, even SFT can generalize with its updates constrained to this direction. This finding offers two practical benefits by combining the strong generalization of on-policy paradigms with the advantages of SFT, including the high training efficiency and ability to leverage high-quality trajectories. For efficiency, we identify update directions that support strong generalization using a few on-policy training steps, and subsequently apply OPSFT to achieve high training efficiency. For leveraging high-quality trajectories, OPSFT can utilize these trajectories to continue improving a post-trained model along its update direction without disrupting the ability learned from on-policy training.