Score
Designs, implements, and evaluates models and scoring functions that assign numeric or ordinal scores to entities to predict risk, quality, priority, propensity, or aesthetic/comparative rankings; this includes statistical and machine‑learning score models, score‑based modeling, and the creation of scoring standards and thresholds. Builds the pipelines, calibration, validation, ranking/threshold logic, and monitoring metrics needed to produce, compare, and maintain reliable, interpretable scores.
This paper addresses the lack of a rigorous theoretical foundation for consistency of scoring functions under variable transformations, specifically examining conditions for consistency and identifiability when predictions and observations undergo one-sided or bijective transformations. Method: We establish formal necessary and sufficient conditions for (strict) consistency and identifiability under general transformations, integrating scoring function theory, Bregman divergence analysis, and techniques from elicitation and identification function characterization for expectation-like functionals. We introduce novel identifiable functionals—including the *g-transformed expectation* and *g-transformed quantile*—and analyze their elicitation properties. Contribution/Results: Our framework provides the first unified theoretical justification for transformed scoring functions in empirical modeling. It enables principled construction of interpretable and verifiable functionals, with broad applicability to probabilistic forecasting and robust regression. The results bridge theoretical statistics and practical model evaluation, ensuring that transformation-based scoring remains both statistically sound and operationally meaningful.
This work addresses the need for uncertainty quantification in ordinal classification within high-stakes domains such as medicine and finance, where errors of varying severity must be rigorously controlled. Existing conformal prediction methods are limited by their choice of nonconformity functions, which often fail to reflect the inherent ordering of classes. To overcome this, the authors propose a novel conformal prediction approach based on the Ranked Probability Score (RPS), introducing RPS as a natural nonconformity measure that captures ordinal risk. This method yields continuous prediction sets centered around the median, avoids greedy search procedures, and maintains model-agnosticism and computational efficiency. It is applicable to both evaluation-based and grouping-based ordinal tasks. Empirical results across multiple image and tabular ordinal datasets demonstrate that the proposed method achieves a superior trade-off between prediction set width and the severity of miscoverage compared to existing approaches.
Existing tabular foundation models are predominantly evaluated using point-estimate metrics such as RMSE and R², which fail to capture tail behavior of predictive distributions and cannot accommodate the asymmetric risk modeling required in high-stakes domains like finance and clinical decision-making. To address this gap, this work proposes ScoringBench—an open-source benchmark that systematically integrates a diverse set of proper scoring rules, including CRPS, CRLS, interval score, and energy score, enabling comprehensive assessment of probabilistic prediction quality alongside traditional metrics. Empirical results demonstrate that model rankings vary substantially depending on the chosen scoring rule, and no single pretraining objective consistently dominates across all criteria, underscoring the necessity of aligning evaluation metrics with application-specific risk characteristics. The benchmark is publicly released with reproducible leaderboards.
Existing performance ranking methods in entity evaluation struggle to simultaneously satisfy application-specific preferences and theoretical rigor. Method: This paper establishes the first axiomatic, verifiable general theory framework for performance ranking. Grounded in probability theory and order theory, it formally defines core concepts—including performance objects, satisfaction, and importance—and introduces a performance order satisfying axioms such as ranking consistency, along with constructive procedures for deriving such orders. It further proposes a novel parameterized family of universal ranking scores that unifies classical metrics (e.g., accuracy, recall, F1-score) and rigorously proves that several widely used metrics—including precision—violate the ranking consistency axiom. Contribution/Results: The framework provides the first mathematically rigorous yet practically flexible foundation for performance evaluation in computer vision and machine learning, explicitly characterizing the validity boundaries and intrinsic limitations of reliable ranking metrics.
This study addresses the lack of comparability in peer review scores across research topics at top machine learning conferences, which undermines fairness in paper acceptance decisions. Analyzing 50,289 submissions to ICLR from 2021 to 2026, this work provides the first systematic evidence that papers from different topics with identical review scores exhibit up to an eightfold difference in acceptance probability. The disparity stems from a fundamental flaw in the measurement design of the scoring system—not from individual reviewer bias or cultural differences in scoring practices. Through rigorous statistical modeling and attribution analysis that controls for multiple confounding factors, the authors propose calibrated review signals and topic-conditioned acceptance rates as core metrics for evaluating fairness, offering critical empirical evidence and policy guidance for reforming conference reviewing mechanisms.
本文评估了不同奖励函数对大语言模型预测性能和行为的影响,比较了五种适当评分规则作为训练目标的效果。
论文解决了LLM评分器因顺序依赖导致决策不一致的问题,提出通过OC-SFT方法训练模型减少顺序影响,提高决策稳定性。
This study addresses the lack of intuitive visualization methods for ordinal regression results, which has hindered their application in fields such as visualization and human-computer interaction. To bridge this gap, the paper proposes, for the first time, the use of modified complementary cumulative distribution function (mCCDF) plots to visualize outputs from cumulative link ordinal regression models. This approach not only fills a critical void in the clear and direct representation of ordinal regression outcomes but also effectively conveys key conclusions consistent with those erroneously derived when treating ordinal variables as continuous. By doing so, the method substantially enhances the interpretability and communicability of model results, offering a principled yet accessible visual framework for practitioners and researchers alike.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
Existing evaluation methods struggle to effectively assess agent performance in subjective, context-dependent, long-horizon enterprise tasks due to their reliance on binary correctness judgments. To address this limitation, this work proposes LH-Bench, a three-pillar evaluation framework that integrates expert-designed rubrics, step-level ground-truth artifact annotations, and pairwise human preference comparisons to enable fine-grained and scalable quantitative assessment. Validation on two real-world scenarios—Figma-to-code translation and procedural content generation—demonstrates that expert-crafted rubrics achieve substantially higher inter-rater agreement than LLM-generated ones (Cohen’s Kappa: 0.60 vs. 0.46), and human preference data significantly align with the framework’s rankings (p < 0.05). The associated dataset has been publicly released.