An Empirical Study of SFT-DPO Interaction and Parameterization in Small Language Models

📅 2026-03-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic investigation into the interplay between supervised fine-tuning (SFT) and direct preference optimization (DPO), as well as parameter-efficient strategies, under conditions of small-scale language models and limited data. Using a GPT-2–sized decoder, the authors systematically compare training paradigms including SFT-only, DPO-only, and SFT followed by DPO, evaluating each with both full-parameter fine-tuning (FFT) and LoRA. Results demonstrate that FFT consistently and significantly outperforms LoRA, which fails to deliver practical speedups. Notably, when DPO’s preference construction aligns with the supervised objective, DPO alone—without SFT pretraining—achieves comparable performance, challenging the conventional assumption that SFT must precede preference-based alignment. The findings indicate that, in small-model settings, full-parameter SFT remains the dominant factor for achieving strong performance.

Technology Category

Machine Learning: Learning Preferences or RankingsSearch and Optimization: Learning to SearchNatural Language Processing: Safety and Robustness

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Direct Preference Optimization (DPO) is widely used after supervised fine-tuning (SFT) to align language models, yet empirical behavior under small backbones and modest data is under-specified. We systematically compare SFT-only, DPO-only, and staged SFT-to-DPO training alongside full fine-tuning (FFT) versus LoRA on a GPT-2-scale decoder, evaluating paraphrase detection and Shakespearean sonnet continuation. DPO yields small, task-dependent gains over strong SFT and can match competitive SFT accuracy without a warm start when the preference construction closely parallels the supervised objective. In contrast, parameterization dominates: FFT consistently outperforms LoRA at matched training depth, and LoRA does not reduce wall-clock time on our hardware. These findings indicate that, in this small-scale regime, supervised full-parameter adaptation remains the primary performance lever, while preference optimization and low-rank adaptation provide limited marginal returns.
Problem

Research questions and friction points this paper is trying to address.

Supervised Fine-Tuning
Direct Preference Optimization
Small Language Models
Parameterization
Empirical Study
Innovation

Methods, ideas, or system contributions that make the work stand out.

Direct Preference Optimization
Supervised Fine-Tuning
LoRA
Full Fine-Tuning
Small Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuming Feng
Department of Computer Science, Stanford University
C
Christy Yang
Department of Computer Science, Stanford University