On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalization capability of Supervised Fine-Tuning (SFT) and its neglect of parameter update directions by proposing OPSFT. The core innovation lies in, for the first time, reformulating on-policy cumulative update directions as explicit optimization constraints for SFT. This reveals how dynamically adjusting update directions facilitates generalization, thereby enabling cross-paradigm transfer of generalization capabilities. By integrating the computational efficiency of SFT with the high-quality trajectory advantages of on-policy methods, OPSFT unifies both paradigms. Experimental results demonstrate that OPSFT significantly enhances both the generalization performance and training efficiency of SFT. Furthermore, it supports the continuous improvement of large language model capabilities through the incorporation of high-quality data.
📝 Abstract
The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for improving the generalization of other paradigms such as supervised fine-tuning (SFT). To address this limitation, we investigate whether there exists a specific on-policy update behavior that can achieve such improvements. First, our analyses reveal that SFT updates parameters along consistent directions, while the on-policy paradigm continuously adjusts the direction during training. This difference inspires us to focus on the cumulative update direction of each parameter as a promising behavior. Then, we evaluate its effectiveness for improving generalization by proposing On-Policy direction-constrained Supervised Fine-Tuning (OPSFT), which constrains SFT updates to the direction identified by on-policy paradigms. The strong performance of OPSFT indicates that the generalization advantage of on-policy paradigms can be transferred to SFT through the parameter update direction. Once such a direction is identified, even SFT can generalize with its updates constrained to this direction. This finding offers two practical benefits by combining the strong generalization of on-policy paradigms with the advantages of SFT, including the high training efficiency and ability to leverage high-quality trajectories. For efficiency, we identify update directions that support strong generalization using a few on-policy training steps, and subsequently apply OPSFT to achieve high training efficiency. For leveraging high-quality trajectories, OPSFT can utilize these trajectories to continue improving a post-trained model along its update direction without disrupting the ability learned from on-policy training.
Problem

Research questions and friction points this paper is trying to address.

On-Policy Post-Training
Supervised Fine-Tuning
Generalization
Parameter Update Direction
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Post-Training
Supervised Fine-Tuning
Parameter Update Direction
Generalization
OPSFT
S
Shufan Shen
State Key Lab. of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Z
Zhongni Hou
Meituan
J
Junshu Sun
State Key Lab. of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Y
Yufei Zhang
Meituan
W
Wei Lin
Meituan
Guojun Yin
Guojun Yin
Meituan, University of Science and Technology of China
MultimodalityComputer VisionFoundation ModelsDeep LearningImage/Video Processing
Qingming Huang
Qingming Huang
University of the Chinese Academy of Sciences
Multimedia Analysis and RetrievalImage and Video ProcessingPattern RecognitionComputer VisionVideo Coding
S
Shuhui Wang
State Key Lab. of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences