Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process

📅 2024-05-20
🏛️ arXiv.org
📈 Citations: 5
✨ Influential: 0
📄 PDF
🤖 AI Summary
A paradigmatic gap exists between supervised fine-tuning (SFT) and preference optimization (PO), hindering unified alignment. Method: We propose intuitive fine-tuning (IFT), a single-stage, single-strategy framework that unifies SFT and PO within an MDP formulation—modeling alignment as token-level preference estimation coupled with transition optimization. We theoretically show that SFT is a degenerate case of PO under zero preference signals, and introduce temporal residual connections to enable end-to-end alignment using only non-preference-labeled data at SFT-scale volume. Contribution/Results: Experiments demonstrate that IFT matches or surpasses the performance of standard two-stage SFT+PO across generation, reasoning, and fact-following tasks. Interpretable analysis on Frozen Lake further validates its policy efficacy. To our knowledge, this is the first work to achieve a unified optimization paradigm for SFT and PO.

Technology Category

Machine Learning: Imitation Learning & Inverse Reinforcement LearningSearch and Optimization: Learning to SearchHumans and AI: Learning Human Values and Preferences

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Supervised Fine-Tuning (SFT) and Preference Optimization (PO) are two fundamental processes for enhancing the capabilities of Language Models (LMs) post pre-training, aligning them better with human preferences. Although SFT advances in training efficiency, PO delivers better alignment, thus they are often combined. However, common practices simply apply them sequentially without integrating their optimization objectives, ignoring the opportunities to bridge their paradigm gap and take the strengths from both. To obtain a unified understanding, we interpret SFT and PO with two sub-processes -- Preference Estimation and Transition Optimization -- defined at token level within the Markov Decision Process (MDP) framework. This modeling shows that SFT is only a specialized case of PO with inferior estimation and optimization. PO evaluates the quality of model's entire generated answer, whereas SFT only scores predicted tokens based on preceding tokens from target answers. Therefore, SFT overestimates the ability of model, leading to inferior optimization. Building on this view, we introduce Intuitive Fine-Tuning (IFT) to integrate SFT and Preference Optimization into a single process. IFT captures LMs' intuitive sense of the entire answers through a temporal residual connection, but it solely relies on a single policy and the same volume of non-preference-labeled data as SFT. Our experiments show that IFT performs comparably or even superiorly to sequential recipes of SFT and some typical Preference Optimization methods across several tasks, particularly those requires generation, reasoning, and fact-following abilities. An explainable Frozen Lake game further validates the effectiveness of IFT for getting competitive policy.
Problem

Research questions and friction points this paper is trying to address.

Bridging gap between SFT and PO alignment methods
Improving token-level preference estimation and optimization
Unifying SFT and PO into a single efficient process
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates SFT and PO into single process
Uses token-level MDP for optimization
Relies on single policy and non-preference data
🔎 Similar Papers
2024-06-05arXiv.orgCitations: 1
💼 Related Jobs
No related jobs found.
Tsinghua University | Frontis.AI
Ermo Hua
Ermo Hua
Tsinghua University
Physics-driven Foundation Model
B
Biqing Qi
Department of Electronic Engineering, Tsinghua University, Beijing, China; Frontis.AI, Beijing, China
Kaiyan Zhang
Kaiyan Zhang
Tsinghua University
Foundation ModelCollective IntelligenceScientific Intelligence
Y
Yue Yu
Department of Electronic Engineering, Tsinghua University, Beijing, China
N
Ning Ding
Department of Electronic Engineering, Tsinghua University, Beijing, China; Frontis.AI, Beijing, China
Xingtai Lv
Xingtai Lv
Tsinghua University
Large Language ModelNatural Language Processing
K
Kai Tian
Department of Electronic Engineering, Tsinghua University, Beijing, China; Frontis.AI, Beijing, China
B
Bowen Zhou
Department of Electronic Engineering, Tsinghua University, Beijing, China