Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work reveals a theoretical unification between supervised fine-tuning (SFT) and preference learning (e.g., DPO) within the optimal policy–reward subspace, showing that SFT implicitly performs reward learning while its conventional KL-divergence regularization fails to constrain model updates due to optimization dynamics. To address this, we propose three innovations: (1) theoretically, we reformulate the SFT objective using f-divergence, broadening the theoretical validity of the logits–Q-function correspondence for large language models; (2) algorithmically, we introduce a learning-rate decay schedule to restore the constraining effect of the KL term; and (3) empirically, we demonstrate up to 25% relative improvement and a 6 percentage-point absolute increase in win rate on instruction-following benchmarks, significantly enhancing post-DPO model performance.

Technology Category

Search and Optimization: Learning to SearchMachine Learning: Learning Preferences or RankingsHumans and AI: Learning Human Values and Preferences

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Post-training processes are essential phases in grounding pre-trained language models to real-world tasks, with learning from demonstrations or preference signals playing a crucial role in this adaptation. We present a unified theoretical framework bridging Supervised Fine-Tuning (SFT) and preference learning in Large Language Model (LLM) post-training. Through rigorous mathematical derivation, we demonstrate that both SFT and preference learning methods like Direct Preference Optimization (DPO) operate within the same optimal policy-reward subspace, with SFT representing a special case of implicit reward learning. Our analysis reveals a critical limitation in conventional SFT: the KL divergence term in distribution matching becomes constant with respect to the policy during optimization, failing to constrain model updates. To address this, we propose a simple yet effective learning rate reduction approach that yields significant performance improvements (up to extbf{25%} relative gain and extbf{6%} absolute win rate increase in instruction following tasks. Additionally, we derive alternative SFT objectives from various f-divergence functions that preserve the KL term during optimization, further enhancing post-DPO model performance. Finally, we extend the theoretical relationship between LLM logits and Q-functions from preference learning to the SFT context, providing mathematical derivations and experimental validation.
Problem

Research questions and friction points this paper is trying to address.

Unify SFT and DPO in a theoretical framework.
Address SFT's KL divergence limitation in optimization.
Extend logit-Q-function theory to SFT context.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified framework connects SFT and DPO via implicit rewards
Learning rate reduction boosts performance significantly
Alternative SFT objectives enhance post-DPO model performance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Bo Wang
School of Computer Science, Fudan University
Q
Qinyuan Cheng
School of Computer Science, Fudan University
R
Runyu Peng
School of Computer Science, Fudan University
Rong Bao
Rong Bao
PhD student, Fudan University
AlignmentGenerative AIReinforcement Learning
Peiji Li
Peiji Li
Fudan University
Qipeng Guo
Qipeng Guo
Fudan University
L
Linyang Li
Shanghai Artificial Intelligence Laboratory
Zhiyuan Zeng
Zhiyuan Zeng
Paul G. Allen School of Computer Science & Engineering, University of Washington
Natural Language ProcessingLarge Language Models
Yunhua Zhou
Yunhua Zhou
Fudan University
Machine LearningNatural Language Processing
X
Xipeng Qiu
School of Computer Science, Fudan University