Score
Designs and builds models, loss functions, and training pipelines that learn, represent, and optimize preferences expressed as pairwise or ordinal comparisons and explicit preference labels, using pairwise ranking losses, direct preference optimization, and preference-based fine-tuning or post-training. This competence includes methods for eliciting and modeling preferences (pairwise elicitation, social or temporal preferences, confidence-aware and unbiased estimators, interactive or dual-negative strategies), calibrating evaluator scores, and integrating preference objectives into generators, retrievers, or other decision modules.
Preference learning lacks a mature theoretical foundation for evaluation. Method: This paper proposes a unified theoretical framework centered on win rate—the probability that a model prefers the correct response—and rigorously proves it is the unique metric satisfying both preference-relation consistency and rationality under data-distribution priors. Based on this, mainstream methods are systematically categorized into win-rate optimization (WRO) and non-WRO classes. Contribution/Results: We establish that WRO methods enjoy dual guarantees of statistical consistency and optimization tractability, whereas canonical non-WRO approaches—including DPO and SFT—exhibit fundamental theoretical limitations. Further analysis reveals that practical model performance is predominantly constrained by optimization difficulty rather than objective design; thus, optimization success rate is a stronger predictor of empirical performance than objective choice. These results provide a verifiable, principled foundation for evaluation, diagnostic analysis, and algorithm design in preference learning.
This work addresses the fragmented landscape of preference learning in large language models, where numerous methods exist without a unifying theoretical foundation, hindering principled practice. We propose the first unified triaxial framework that decomposes existing approaches—such as RLHF, DPO, IPO, KTO, and SimPO—into three orthogonal dimensions: preference modeling, regularization mechanisms, and data distribution. Through theoretical modeling, formal proofs, and extensive empirical validation across more than 50 studies, we delineate the theoretical boundaries between online and offline learning, derive scaling laws governing reward over-optimization, and synthesize actionable guidelines for practitioners. This effort advances preference learning from an empirically driven paradigm toward a theoretically grounded discipline.
This work addresses a key limitation of existing Direct Preference Optimization (DPO) methods, which rely solely on pairwise preference signals and neglect the quantitative differences in response quality, leading to ambiguous training signals and suboptimal optimization efficiency. To overcome this, we propose a novel preference optimization algorithm that, for the first time, incorporates explicit score gaps into the DPO framework. Our approach preserves the advantage of not requiring an explicit reward model while leveraging fine-grained relative quality information to enhance alignment. By designing a loss function that accounts for score differences, the method enjoys faster theoretical statistical convergence and demonstrates robustness to scoring noise. Extensive experiments show consistent and significant improvements over current DPO variants across multiple large language models and evaluation benchmarks, with stable performance gains even when provided with inaccurate scores.
Existing Bradley–Terry–based alignment methods (e.g., DPO, RLHF) rely on pairwise comparisons, which inadequately capture the full ranking structure among multiple responses and thus limit alignment with diverse human preferences. This work introduces, for the first time, the information retrieval metric Normalized Discounted Cumulative Gain (NDCG) into LLM alignment, proposing a differentiable ordinal preference optimization framework. We construct a surrogate loss via a differentiable NDCG approximation to enable end-to-end multi-response ranking optimization. Furthermore, we integrate ordinal preference modeling with a negative-sample pool expansion strategy to mitigate interference from trivial negatives. Extensive evaluation on benchmarks—including AlpacaEval—demonstrates significant improvements over DPO and RLHF, validating the method’s effectiveness in enhancing both response quality and fidelity to human ordinal preferences.
Large language models (LLMs) often generate harmful outputs misaligned with human preferences, and existing preference alignment methods lack rigorous generalization guarantees. Method: This work investigates the generalization capability of direct preference optimization (DPO) under diverse human values. We establish the first theoretical framework for DPO generalization error under finite-step gradient updates, deriving tight generalization bounds via reward-margin trajectory analysis. Our approach integrates theoretical analysis—combining generalization bounds with reward-margin modeling—with empirical validation across multiple mainstream LLMs. Contribution/Results: We prove that DPO-trained models correctly rank preferred responses over dispreferred ones on unseen samples with high probability. Moreover, generalization performance improves robustly with increasing value diversity and sample size. This work bridges a critical gap in DPO theory, providing the first verifiable, generalization-theoretic foundation for value-aligned LLM training.
When real-world preference data violate the assumptions of the Bradley–Terry (BT) model, it remains unclear what the BT learning procedure actually recovers. This work addresses this gap by formalizing preference information through the conditional preference response distribution (CPRD) derived from triplet comparison data and establishes, for the first time, a data-centric theoretical framework to characterize the target recovered by the BT model under non-ideal conditions. By integrating graph connectivity, statistical learning theory, and BT model analysis, we precisely delineate the conditions under which the BT model is valid and reveal the critical roles of the marginal distribution and preference graph connectivity in determining sample efficiency. Our results provide a rigorous theoretical foundation for preference learning and alignment tasks under general preference data settings.
This study investigates the optimal timing for incorporating preference embeddings in artificial intelligence systems—whether during training or as a post-processing step. Drawing on information design theory, the work proposes a unified welfare framework that is agnostic to specific decision objectives, revealing how preference embedding induces a contraction in posterior means and thereby affects the value of information. By integrating convexity analysis of information value with models of cognitive constraints, the paper demonstrates that, in the absence of cognitive frictions, preference-agnostic training weakly dominates preference-embedded approaches. However, under human cognitive limitations, embedding preferences during training enhances performance through automatic threshold computation, offering theoretical justification for modular AI architectures.
This study investigates whether sequential preference optimization leads to uniform forgetting of previously learned preferences and how this phenomenon is influenced by the relationships among preference objectives. Using Llama-3.1-8B-Instruct with LoRA adapters, the authors apply Direct Preference Optimization (DPO) sequentially across four distinct preference settings, employing length-normalized margins and quartile-based decomposition for fine-grained analysis alongside gradient diagnostics. The findings reveal that sequential DPO does not induce uniform forgetting; instead, it exhibits diverse behaviors ranging from degradation and stability to positive transfer. Objective compatibility and signal strength emerge as key determinants of these dynamics. High-confidence preference pairs can either improve or deteriorate across stages, and inter-stage gradients are nearly orthogonal, suggesting that gradient interference is not the primary cause of forgetting. These insights offer new design principles for multi-objective alignment.
This work addresses the critical challenge of accurately modeling preference functions that aggregate multidimensional criteria into holistic judgments in settings such as admissions and medical diagnosis. Departing from conventional assumptions of linearity or strong structural forms, the paper proposes the first robust nonparametric learning algorithm that achieves optimal performance without requiring any prior knowledge of the preference structure, assuming only monotonic non-decreasing behavior across each criterion. Theoretical analysis demonstrates the severe consequences of common model misspecifications, while experiments on both synthetic and real-world data confirm that the method maintains statistical efficiency under linear preferences and reliably recovers true evaluator preferences in general cases. Notably, the approach effectively uncovers key behavioral differences between human evaluators and large language models in their assessment strategies.
This work addresses the challenge of efficiently selecting the most informative pairwise comparisons under limited annotation budgets to improve alignment in preference-based large language model post-training. Framing comparison selection as a sampling design problem within the Direct Preference Optimization (DPO) framework, this study establishes the first theoretical connection between comparison pair sampling and policy suboptimality, deriving matching upper and lower bounds. Building on this analysis, the authors propose an explicit sampling criterion based on the Fisher information matrix to guide data acquisition. Experimental results demonstrate that the proposed method significantly outperforms existing heuristic strategies on both synthetic benchmarks and real-world language model post-training tasks, achieving substantially higher sample efficiency.