🤖 AI Summary
This work reveals a theoretical unification between supervised fine-tuning (SFT) and preference learning (e.g., DPO) within the optimal policy–reward subspace, showing that SFT implicitly performs reward learning while its conventional KL-divergence regularization fails to constrain model updates due to optimization dynamics. To address this, we propose three innovations: (1) theoretically, we reformulate the SFT objective using f-divergence, broadening the theoretical validity of the logits–Q-function correspondence for large language models; (2) algorithmically, we introduce a learning-rate decay schedule to restore the constraining effect of the KL term; and (3) empirically, we demonstrate up to 25% relative improvement and a 6 percentage-point absolute increase in win rate on instruction-following benchmarks, significantly enhancing post-DPO model performance.
📝 Abstract
Post-training processes are essential phases in grounding pre-trained language models to real-world tasks, with learning from demonstrations or preference signals playing a crucial role in this adaptation. We present a unified theoretical framework bridging Supervised Fine-Tuning (SFT) and preference learning in Large Language Model (LLM) post-training. Through rigorous mathematical derivation, we demonstrate that both SFT and preference learning methods like Direct Preference Optimization (DPO) operate within the same optimal policy-reward subspace, with SFT representing a special case of implicit reward learning. Our analysis reveals a critical limitation in conventional SFT: the KL divergence term in distribution matching becomes constant with respect to the policy during optimization, failing to constrain model updates. To address this, we propose a simple yet effective learning rate reduction approach that yields significant performance improvements (up to extbf{25%} relative gain and extbf{6%} absolute win rate increase in instruction following tasks. Additionally, we derive alternative SFT objectives from various f-divergence functions that preserve the KL term during optimization, further enhancing post-DPO model performance. Finally, we extend the theoretical relationship between LLM logits and Q-functions from preference learning to the SFT context, providing mathematical derivations and experimental validation.