🤖 AI Summary
This study addresses the vulnerability of Direct Preference Optimization (DPO) to absorbing and amplifying systematic biases—such as gender and length preferences—inherent in human annotations, which compromises model fairness. To mitigate this, we propose BA-DPO, a generalized bias-adjustment framework that introduces attribute-specific bias parameters for individual annotators. Through convex optimization, this approach eliminates arbitrary attribute biases and supports specifying target attribute rates to achieve statistical parity. We establish the convexity of the objective function, marking the first generalizable framework of its kind. Empirical evaluations demonstrate that BA-DPO mitigates 81%–95% of bias drift and substantially suppresses response length inflation. Furthermore, it preserves generation quality while maintaining a KL divergence no greater than that of standard DPO.
📝 Abstract
Preference-based alignment methods such as Direct Preference Optimization (DPO) use pairwise preferences labeled by human annotators to fine-tune language models. However, annotators carry systematic biases toward some attributes: a name that signals a gender or an ethnicity, a persona, a language variety, a formatting convention, or length. If not properly addressed, these systematic biases can be absorbed and amplified during alignment. Existing methods address length bias or annotator disagreement, but fail to eliminate biases toward arbitrary attributes. To address this limitation, we propose Bias-Adjusted DPO (BA-DPO), a generalization of DPO that adds one bias parameter per annotator toward responses carrying a declared attribute. We prove that the objective is convex in the bias parameters and that the votes identify each annotator's bias up to a shared constant. The remaining constant is what fixes the aligned model's attribute rate: by default the rate of the reference model, or a target rate, which we use to bring a biased policy to statistical parity. On a corpus with planted biases, DPO drives the attribute from a balanced start to probability 0.96 and BA-DPO removes 81 to 95\% of that shift; on MultiPref with real annotators it removes about half of DPO's lengthening. Both hold at 0.5B with full fine-tuning and at 8B with LoRA, at no higher KL than DPO and no loss in judged quality.