Boosting Direct Preference Optimization with Penalization

πŸ“… 2026-06-10
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses a limitation in existing offline preference optimization methods, such as Direct Preference Optimization (DPO), which utilize only the chosen and rejected responses from static datasets while neglecting the greedy response generated by the reference model for the same prompt as a potential supervisory signal. The authors propose DPOP, which extends the DPO framework by introducing a gated penalty term that suppresses the reference model’s greedy response only when the policy model assigns lower likelihood to the preferred response than to the rejected one. Combined with length normalization to enhance fairness, this approach uniquely leverages the reference model’s own greedy output as a conditionally activated supervision signal, substantially improving preference learning. On AlpacaEval 2.0, DPOP achieves length-controlled win rate improvements of 5.3% and 4.4% over baselines using Llama-3-8b-instruct and Gemma-2-9b-instruct, respectively.
πŸ“ Abstract
Offline preference optimization has become a practical substitute for reinforcement learning from human feedback, but pairwise objectives such as Direct Preference Optimization (DPO) and its variants use only the chosen and rejected responses stored in a static dataset. This leaves a useful signal unused: the response that the reference model itself would generate for the same prompt. We propose Direct Preference Optimization with Penalization (DPOP), a simple extension of DPO that augments the base preference loss with a gated penalty on reference-greedy responses. DPOP activates this penalty only when the current policy still assigns a lower likelihood to the preferred response than to the rejected response. On AlpacaEval 2.0, DPOP improves length-controlled win rate over DPO, SimPO, and AlphaDPO on both Llama-3-8b-it and Gemma-2-9b-it, achieving relative gains of 5.3\% and 4.4\% over baselines on the two models, respectively. Ablations further show that a SimNPO-style length-normalized penalty is stronger than NPO and token-level unlikelihood in this setting.
Problem

Research questions and friction points this paper is trying to address.

Offline Preference Optimization
Direct Preference Optimization
Reference Model
Preference Learning
Static Dataset
Innovation

Methods, ideas, or system contributions that make the work stand out.

Direct Preference Optimization
Penalization
Offline Preference Learning
Reference-Greedy Response
Length-Normalized Penalty
P
Pengwei Sun
Stanford University