Score
Designs, trains, and evaluates models and methods that produce ordered outputs or continuous scalar scores for items by optimizing ranking-specific objectives (pairwise, listwise, probabilistic) to enforce relative orderings. Builds ranking algorithms and loss functions that improve discrimination among items or classes, including across distant or imbalanced categories.
Offline model-based optimization (MBO) faces a fundamental challenge: regression models trained on static datasets suffer from out-of-distribution errors, leading to overestimation of suboptimal designs and misguiding the optimization process. This work observes that MBO’s core objective is to **identify promising design rankings**, not to predict absolute performance scores accurately. Accordingly, we propose the first integration of **Learning to Rank (LTR)** into offline MBO. Instead of minimizing mean squared error, our method employs pairwise or listwise ranking losses within an offline reinforcement learning framework to explicitly model relative design preferences. We further derive a theoretical upper bound on the generalization error of ranking loss in this setting. Evaluated across diverse benchmark tasks, our approach consistently outperforms 20 state-of-the-art methods, achieving superior robustness and higher-quality optimal solutions.
Existing performance ranking methods in entity evaluation struggle to simultaneously satisfy application-specific preferences and theoretical rigor. Method: This paper establishes the first axiomatic, verifiable general theory framework for performance ranking. Grounded in probability theory and order theory, it formally defines core concepts—including performance objects, satisfaction, and importance—and introduces a performance order satisfying axioms such as ranking consistency, along with constructive procedures for deriving such orders. It further proposes a novel parameterized family of universal ranking scores that unifies classical metrics (e.g., accuracy, recall, F1-score) and rigorously proves that several widely used metrics—including precision—violate the ranking consistency axiom. Contribution/Results: The framework provides the first mathematically rigorous yet practically flexible foundation for performance evaluation in computer vision and machine learning, explicitly characterizing the validity boundaries and intrinsic limitations of reliable ranking metrics.
This paper addresses the problem of automatically synthesizing concise linear scoring functions from given relations and tuple rankings, without prior knowledge of ranking functions. The proposed method introduces Symbolic Gradient Descent (Sym-GD), an approximation algorithm, and formulates the synthesis task as a Mixed-Integer Linear Program (MILP) that supports position-aware error minimization and customizable weight constraints. To overcome the limitations of traditional polynomial-time algorithms—which solve subproblems in isolation—the approach leverages LP decomposition and constrained optimization techniques. Experimental results demonstrate that the method achieves speedups of several orders of magnitude over state-of-the-art baselines, while significantly improving accuracy and scalability on large-scale real-world datasets. It successfully generates linear scoring functions that are both more accurate and inherently interpretable.
In critical ranking-dependent decision-making scenarios—such as hiring, admissions, and credit scoring—existing explainable AI methods (e.g., SHAP) fail to adapt effectively, as their underlying objective functions (e.g., classification or regression loss) are fundamentally misaligned with ranking-specific goals (e.g., rank position, top-k inclusion, pairwise preferences). To address this gap, we propose ShaRP: the first Shapley-value-based framework for ranking explainability. ShaRP extends Shapley attribution to ranking and preference modeling by introducing ranking-specific utility functions—including rank loss, top-k indicator, and pairwise preference—and deriving corresponding feature importance formulations. It supports both score-based ranking and learning-to-rank paradigms and is compatible with tabular data. Extensive experiments demonstrate that ShaRP significantly outperforms baselines in explanation fidelity, computational efficiency, and task coverage. By providing rigorous, flexible, and verifiable attribution, ShaRP establishes a principled foundation for interpretable ranking decisions.
This work addresses a key limitation of existing Direct Preference Optimization (DPO) methods, which rely solely on pairwise preference signals and neglect the quantitative differences in response quality, leading to ambiguous training signals and suboptimal optimization efficiency. To overcome this, we propose a novel preference optimization algorithm that, for the first time, incorporates explicit score gaps into the DPO framework. Our approach preserves the advantage of not requiring an explicit reward model while leveraging fine-grained relative quality information to enhance alignment. By designing a loss function that accounts for score differences, the method enjoys faster theoretical statistical convergence and demonstrates robustness to scoring noise. Extensive experiments show consistent and significant improvements over current DPO variants across multiple large language models and evaluation benchmarks, with stable performance gains even when provided with inaccurate scores.
This work reframes the evaluation of explanation quality as a learning-to-rank problem, moving beyond conventional approaches that rely on generating a single optimal explanation or pointwise regression, which struggle to distinguish among explanations of varying quality levels. The study introduces listwise ranking methods—specifically ListNet, LambdaRank, and RankNet—to train a reward model capable of performing relative quality assessment over multiple candidate explanations while preserving their ordinal structure. Experimental results demonstrate that ranking-based losses consistently outperform regression-based counterparts across all domains. Furthermore, policy optimization using ranking-derived rewards achieves stable convergence, whereas regression-based rewards fail entirely. The findings also highlight that data quality exerts a more decisive influence than model scale, enabling smaller models to match the performance of significantly larger ones.
This study investigates the robustness of subset rankings under ordinal aggregation by merging similar items in an item similarity graph, assuming additive evaluation metrics. The problem is formulated as four classes of combinatorial optimization tasks, aiming to maximize or minimize either the absolute or relative rank of a given subset. The work provides the first systematic characterization of the computational complexity of ranking optimization with partitioning operations, establishing NP-hardness for most variants while developing exact and approximation algorithms tailored to realistic, structured graph topologies. The proposed methodology is successfully applied to assess the robustness of rankings of greenhouse gas emission sources, demonstrating its practical utility across domains.
This work addresses the limitations of offline model-based optimization, which often relies on regression modeling and suffers from distributional mismatch between training data and near-optimal designs. From a learnability perspective, the authors reformulate the optimization problem as a ranking task that distinguishes high-performing from suboptimal designs, proposing a ranking-centric, optimization-oriented risk framework. The framework identifies distributional mismatch as the primary source of error and theoretically demonstrates the superiority of ranking over conventional regression approaches. By integrating an optimization-aware ranking risk, a unified learning theory, and a distribution-aware algorithm, the method significantly outperforms 20 existing baselines across diverse tasks, validating the efficacy of the ranking perspective and exposing the fundamental limitation of overly optimistic extrapolation in offline optimization.