preference modeling

Designs and builds models, loss functions, and training pipelines that learn, represent, and optimize preferences expressed as pairwise or ordinal comparisons and explicit preference labels, using pairwise ranking losses, direct preference optimization, and preference-based fine-tuning or post-training. This competence includes methods for eliciting and modeling preferences (pairwise elicitation, social or temporal preferences, confidence-aware and unbiased estimators, interactive or dual-negative strategies), calibrating evaluator scores, and integrating preference objectives into generators, retrievers, or other decision modules.

preferencemodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.82
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$217K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Preference learning made easy: Everything should be understood through win rate

Feb 14, 2025
LH
Lily H. Zhang
🏛️ New York University

Preference learning lacks a mature theoretical foundation for evaluation. Method: This paper proposes a unified theoretical framework centered on win rate—the probability that a model prefers the correct response—and rigorously proves it is the unique metric satisfying both preference-relation consistency and rationality under data-distribution priors. Based on this, mainstream methods are systematically categorized into win-rate optimization (WRO) and non-WRO classes. Contribution/Results: We establish that WRO methods enjoy dual guarantees of statistical consistency and optimization tractability, whereas canonical non-WRO approaches—including DPO and SFT—exhibit fundamental theoretical limitations. Further analysis reveals that practical model performance is predominantly constrained by optimization difficulty rather than objective design; thus, optimization success rate is a stronger predictor of empirical performance than objective choice. These results provide a verifiable, principled foundation for evaluation, diagnostic analysis, and algorithm design in preference learning.

Analyzing methods as win rate optimization or non-WRO.Improving optimization of WRO objectives in practice.Understanding preference learning through win rate.

This work addresses the fragmented landscape of preference learning in large language models, where numerous methods exist without a unifying theoretical foundation, hindering principled practice. We propose the first unified triaxial framework that decomposes existing approaches—such as RLHF, DPO, IPO, KTO, and SimPO—into three orthogonal dimensions: preference modeling, regularization mechanisms, and data distribution. Through theoretical modeling, formal proofs, and extensive empirical validation across more than 50 studies, we delineate the theoretical boundaries between online and offline learning, derive scaling laws governing reward over-optimization, and synthesize actionable guidelines for practitioners. This effort advances preference learning from an empirically driven paradigm toward a theoretically grounded discipline.

human alignmentlarge language modelsmethod selection

This work addresses a key limitation of existing Direct Preference Optimization (DPO) methods, which rely solely on pairwise preference signals and neglect the quantitative differences in response quality, leading to ambiguous training signals and suboptimal optimization efficiency. To overcome this, we propose a novel preference optimization algorithm that, for the first time, incorporates explicit score gaps into the DPO framework. Our approach preserves the advantage of not requiring an explicit reward model while leveraging fine-grained relative quality information to enhance alignment. By designing a loss function that accounts for score differences, the method enjoys faster theoretical statistical convergence and demonstrates robustness to scoring noise. Extensive experiments show consistent and significant improvements over current DPO variants across multiple large language models and evaluation benchmarks, with stable performance gains even when provided with inaccurate scores.

Alignment ProblemDirect Preference OptimizationFoundation Models

Ordinal Preference Optimization: Aligning Human Preferences via NDCG

Oct 06, 2024
YZ
Yang Zhao
🏛️ Tsinghua University | University of Michigan, Ann Arbor | University of Florida

Existing Bradley–Terry–based alignment methods (e.g., DPO, RLHF) rely on pairwise comparisons, which inadequately capture the full ranking structure among multiple responses and thus limit alignment with diverse human preferences. This work introduces, for the first time, the information retrieval metric Normalized Discounted Cumulative Gain (NDCG) into LLM alignment, proposing a differentiable ordinal preference optimization framework. We construct a surrogate loss via a differentiable NDCG approximation to enable end-to-end multi-response ranking optimization. Furthermore, we integrate ordinal preference modeling with a negative-sample pool expansion strategy to mitigate interference from trivial negatives. Extensive evaluation on benchmarks—including AlpacaEval—demonstrates significant improvements over DPO and RLHF, validating the method’s effectiveness in enhancing both response quality and fidelity to human ordinal preferences.

Addressing limitations of pairwise methods in listwise ranking tasksAligning LLMs with human preferences for controllable behaviorsImproving ranking accuracy using NDCG-based alignment approach

On the Generalization of Preference Learning with DPO

Aug 06, 2024
SI
Shawn Im
🏛️ University of Wisconsin-Madison

Large language models (LLMs) often generate harmful outputs misaligned with human preferences, and existing preference alignment methods lack rigorous generalization guarantees. Method: This work investigates the generalization capability of direct preference optimization (DPO) under diverse human values. We establish the first theoretical framework for DPO generalization error under finite-step gradient updates, deriving tight generalization bounds via reward-margin trajectory analysis. Our approach integrates theoretical analysis—combining generalization bounds with reward-margin modeling—with empirical validation across multiple mainstream LLMs. Contribution/Results: We prove that DPO-trained models correctly rank preferred responses over dispreferred ones on unseen samples with high probability. Moreover, generalization performance improves robustly with increasing value diversity and sample size. This work bridges a critical gap in DPO theory, providing the first verifiable, generalization-theoretic foundation for value-aligned LLM training.

Analyzing generalization scaling with value diversity in preference optimizationAssessing how models generalize after finite gradient steps in trainingProviding generalization error bounds for learning diverse human values

Latest Papers

What's happening recently
View more

When real-world preference data violate the assumptions of the Bradley–Terry (BT) model, it remains unclear what the BT learning procedure actually recovers. This work addresses this gap by formalizing preference information through the conditional preference response distribution (CPRD) derived from triplet comparison data and establishes, for the first time, a data-centric theoretical framework to characterize the target recovered by the BT model under non-ideal conditions. By integrating graph connectivity, statistical learning theory, and BT model analysis, we precisely delineate the conditions under which the BT model is valid and reveal the critical roles of the marginal distribution and preference graph connectivity in determining sample efficiency. Our results provide a rigorous theoretical foundation for preference learning and alignment tasks under general preference data settings.

Bradley-Terry modelconditional preference distributionmodel misspecification

This study investigates the optimal timing for incorporating preference embeddings in artificial intelligence systems—whether during training or as a post-processing step. Drawing on information design theory, the work proposes a unified welfare framework that is agnostic to specific decision objectives, revealing how preference embedding induces a contraction in posterior means and thereby affects the value of information. By integrating convexity analysis of information value with models of cognitive constraints, the paper demonstrates that, in the absence of cognitive frictions, preference-agnostic training weakly dominates preference-embedded approaches. However, under human cognitive limitations, embedding preferences during training enhances performance through automatic threshold computation, offering theoretical justification for modular AI architectures.

decision-making under cognitive constraintsinformation designmodular AI pipelines

This study investigates whether sequential preference optimization leads to uniform forgetting of previously learned preferences and how this phenomenon is influenced by the relationships among preference objectives. Using Llama-3.1-8B-Instruct with LoRA adapters, the authors apply Direct Preference Optimization (DPO) sequentially across four distinct preference settings, employing length-normalized margins and quartile-based decomposition for fine-grained analysis alongside gradient diagnostics. The findings reveal that sequential DPO does not induce uniform forgetting; instead, it exhibits diverse behaviors ranging from degradation and stability to positive transfer. Objective compatibility and signal strength emerge as key determinants of these dynamics. High-confidence preference pairs can either improve or deteriorate across stages, and inter-stage gradients are nearly orthogonal, suggesting that gradient interference is not the primary cause of forgetting. These insights offer new design principles for multi-objective alignment.

Direct Preference Optimizationmulti-objective learningobjective compatibility

This work addresses the critical challenge of accurately modeling preference functions that aggregate multidimensional criteria into holistic judgments in settings such as admissions and medical diagnosis. Departing from conventional assumptions of linearity or strong structural forms, the paper proposes the first robust nonparametric learning algorithm that achieves optimal performance without requiring any prior knowledge of the preference structure, assuming only monotonic non-decreasing behavior across each criterion. Theoretical analysis demonstrates the severe consequences of common model misspecifications, while experiments on both synthetic and real-world data confirm that the method maintains statistical efficiency under linear preferences and reliably recovers true evaluator preferences in general cases. Notably, the approach effectively uncovers key behavioral differences between human evaluators and large language models in their assessment strategies.

evaluator preferencesmodel mismatchmulti-criteria evaluation

This work addresses the challenge of efficiently selecting the most informative pairwise comparisons under limited annotation budgets to improve alignment in preference-based large language model post-training. Framing comparison selection as a sampling design problem within the Direct Preference Optimization (DPO) framework, this study establishes the first theoretical connection between comparison pair sampling and policy suboptimality, deriving matching upper and lower bounds. Building on this analysis, the authors propose an explicit sampling criterion based on the Fisher information matrix to guide data acquisition. Experimental results demonstrate that the proposed method significantly outperforms existing heuristic strategies on both synthetic benchmarks and real-world language model post-training tasks, achieving substantially higher sample efficiency.

comparison selectionlabeling budgetLLM alignment

Hot Scholars

MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
HY

Huaxiu Yao

Assistant Professor of Computer Science and Data Science, UNC Chapel Hill
Machine LearningFoundation ModelsAI AlignmentAI Agent
XH

Xiangnan He

University of Science and Technology of China
RecommendationCausalityBig DataInformation Retrieval
DY

Dawei Yin

Senior Director, Head of Search Science at Baidu
Machine LearningWeb MiningData Mining