🤖 AI Summary
This work addresses the limitation of Direct Preference Optimization (DPO) in modeling the diversity of human preferences. It introduces Mallows ranking theory—previously unexplored in LLM alignment—proposing a Mallows-based preference dispersion index that unifies and generalizes existing preference optimization frameworks. Methodologically, it parameterizes dispersion and reweights preference losses, enabling plug-and-play compatibility with mainstream offline preference optimization algorithms. Empirical evaluation on synthetic bandit, controllable generation, and dialogue tasks demonstrates significant performance gains; integrated as a lightweight plugin into Llama3-Instruct, it improves LC win rate by nearly 2%, confirming strong generalizability and interpretability. Core contributions are threefold: (1) a theoretically grounded, learnable metric for preference dispersion; (2) a unified generalization of DPO; and (3) an efficient, drop-in optimization enhancement for practical deployment.
📝 Abstract
Direct Preference Optimization (DPO) has recently emerged as a popular approach to improve reinforcement learning with human feedback (RLHF), leading to better techniques to fine-tune large language models (LLM). A weakness of DPO, however, lies in its lack of capability to characterize the diversity of human preferences. Inspired by Mallows' theory of preference ranking, we develop in this paper a new approach, the MallowsPO. A distinct feature of this approach is a dispersion index, which reflects the dispersion of human preference to prompts. We show that existing DPO models can be reduced to special cases of this dispersion index, thus unified with MallowsPO. More importantly, we demonstrate (empirically) how to use this dispersion index to enhance the performance of DPO in a broad array of benchmark tasks, from synthetic bandit selection to controllable generations and dialogues, while maintaining great generalization capabilities. MallowsPO is also compatible with other SOTA offline preference optimization methods, boosting nearly 2% extra LC win rate when used as a plugin for fine-tuning Llama3-Instruct.