MallowsPO: Fine-Tune Your LLM with Preference Dispersions

📅 2024-05-23
📈 Citations: 3
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of Direct Preference Optimization (DPO) in modeling the diversity of human preferences. It introduces Mallows ranking theory—previously unexplored in LLM alignment—proposing a Mallows-based preference dispersion index that unifies and generalizes existing preference optimization frameworks. Methodologically, it parameterizes dispersion and reweights preference losses, enabling plug-and-play compatibility with mainstream offline preference optimization algorithms. Empirical evaluation on synthetic bandit, controllable generation, and dialogue tasks demonstrates significant performance gains; integrated as a lightweight plugin into Llama3-Instruct, it improves LC win rate by nearly 2%, confirming strong generalizability and interpretability. Core contributions are threefold: (1) a theoretically grounded, learnable metric for preference dispersion; (2) a unified generalization of DPO; and (3) an efficient, drop-in optimization enhancement for practical deployment.

Technology Category

Machine Learning: Learning Preferences or RankingsHumans and AI: Learning Human Values and PreferencesSearch and Optimization: Learning to Search

Application Category

User Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Direct Preference Optimization (DPO) has recently emerged as a popular approach to improve reinforcement learning with human feedback (RLHF), leading to better techniques to fine-tune large language models (LLM). A weakness of DPO, however, lies in its lack of capability to characterize the diversity of human preferences. Inspired by Mallows' theory of preference ranking, we develop in this paper a new approach, the MallowsPO. A distinct feature of this approach is a dispersion index, which reflects the dispersion of human preference to prompts. We show that existing DPO models can be reduced to special cases of this dispersion index, thus unified with MallowsPO. More importantly, we demonstrate (empirically) how to use this dispersion index to enhance the performance of DPO in a broad array of benchmark tasks, from synthetic bandit selection to controllable generations and dialogues, while maintaining great generalization capabilities. MallowsPO is also compatible with other SOTA offline preference optimization methods, boosting nearly 2% extra LC win rate when used as a plugin for fine-tuning Llama3-Instruct.
Problem

Research questions and friction points this paper is trying to address.

Addresses DPO's inability to capture human preference diversity.
Introduces MallowsPO with a dispersion index for preference ranking.
Enhances DPO performance across various tasks and benchmarks.
Innovation

Methods, ideas, or system contributions that make the work stand out.

MallowsPO introduces a dispersion index for preferences.
Unifies DPO models under Mallows' preference ranking theory.
Enhances DPO performance across diverse benchmark tasks.
🔎 Similar Papers
Columbia University
H
Haoxian Chen
Department of Industrial Engineering and Operations Research, Columbia University
H
Hanyang Zhao
Department of Industrial Engineering and Operations Research, Columbia University
H
Henry Lam
Department of Industrial Engineering and Operations Research, Columbia University
D
David D. Yao
Department of Industrial Engineering and Operations Research, Columbia University
Wenpin Tang
Wenpin Tang
Assistant Professor, Columbia University
Probability TheoryStochastic ProcessesStatisticsMachine Learning