🤖 AI Summary
This study addresses the lack of systematic investigation into the interplay between supervised fine-tuning (SFT) and direct preference optimization (DPO), as well as parameter-efficient strategies, under conditions of small-scale language models and limited data. Using a GPT-2–sized decoder, the authors systematically compare training paradigms including SFT-only, DPO-only, and SFT followed by DPO, evaluating each with both full-parameter fine-tuning (FFT) and LoRA. Results demonstrate that FFT consistently and significantly outperforms LoRA, which fails to deliver practical speedups. Notably, when DPO’s preference construction aligns with the supervised objective, DPO alone—without SFT pretraining—achieves comparable performance, challenging the conventional assumption that SFT must precede preference-based alignment. The findings indicate that, in small-model settings, full-parameter SFT remains the dominant factor for achieving strong performance.
📝 Abstract
Direct Preference Optimization (DPO) is widely used after supervised fine-tuning (SFT) to align language models, yet empirical behavior under small backbones and modest data is under-specified. We systematically compare SFT-only, DPO-only, and staged SFT-to-DPO training alongside full fine-tuning (FFT) versus LoRA on a GPT-2-scale decoder, evaluating paraphrase detection and Shakespearean sonnet continuation. DPO yields small, task-dependent gains over strong SFT and can match competitive SFT accuracy without a warm start when the preference construction closely parallels the supervised objective. In contrast, parameterization dominates: FFT consistently outperforms LoRA at matched training depth, and LoRA does not reduce wall-clock time on our hardware. These findings indicate that, in this small-scale regime, supervised full-parameter adaptation remains the primary performance lever, while preference optimization and low-rank adaptation provide limited marginal returns.