scoring model development

Designs, implements, and evaluates models and scoring functions that assign numeric or ordinal scores to entities to predict risk, quality, priority, propensity, or aesthetic/comparative rankings; this includes statistical and machine‑learning score models, score‑based modeling, and the creation of scoring standards and thresholds. Builds the pipelines, calibration, validation, ranking/threshold logic, and monitoring metrics needed to produce, compare, and maintain reliable, interpretable scores.

scoringmodeldevelopment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.52
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$198K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Transformations of predictions and realizations in consistent scoring functions

Feb 23, 2025
HT
Hristos Tyralis
🏛️ Hellenic Air Force | University of Padova

This paper addresses the lack of a rigorous theoretical foundation for consistency of scoring functions under variable transformations, specifically examining conditions for consistency and identifiability when predictions and observations undergo one-sided or bijective transformations. Method: We establish formal necessary and sufficient conditions for (strict) consistency and identifiability under general transformations, integrating scoring function theory, Bregman divergence analysis, and techniques from elicitation and identification function characterization for expectation-like functionals. We introduce novel identifiable functionals—including the *g-transformed expectation* and *g-transformed quantile*—and analyze their elicitation properties. Contribution/Results: Our framework provides the first unified theoretical justification for transformed scoring functions in empirical modeling. It enables principled construction of interpretable and verifiable functionals, with broad applicability to probabilistic forecasting and robust regression. The results bridge theoretical statistics and practical model evaluation, ensuring that transformation-based scoring remains both statistically sound and operationally meaningful.

Analyzing transformations in realization and prediction variables.Characterizing transformed scoring functions' consistency.Developing novel elicitable functionals for predictive tasks.

This work addresses the need for uncertainty quantification in ordinal classification within high-stakes domains such as medicine and finance, where errors of varying severity must be rigorously controlled. Existing conformal prediction methods are limited by their choice of nonconformity functions, which often fail to reflect the inherent ordering of classes. To overcome this, the authors propose a novel conformal prediction approach based on the Ranked Probability Score (RPS), introducing RPS as a natural nonconformity measure that captures ordinal risk. This method yields continuous prediction sets centered around the median, avoids greedy search procedures, and maintains model-agnosticism and computational efficiency. It is applicable to both evaluation-based and grouping-based ordinal tasks. Empirical results across multiple image and tabular ordinal datasets demonstrate that the proposed method achieves a superior trade-off between prediction set width and the severity of miscoverage compared to existing approaches.

conformal predictionordinal classificationprediction sets

Existing tabular foundation models are predominantly evaluated using point-estimate metrics such as RMSE and R², which fail to capture tail behavior of predictive distributions and cannot accommodate the asymmetric risk modeling required in high-stakes domains like finance and clinical decision-making. To address this gap, this work proposes ScoringBench—an open-source benchmark that systematically integrates a diverse set of proper scoring rules, including CRPS, CRLS, interval score, and energy score, enabling comprehensive assessment of probabilistic prediction quality alongside traditional metrics. Empirical results demonstrate that model rankings vary substantially depending on the chosen scoring rule, and no single pretraining objective consistently dominates across all criteria, underscoring the necessity of aligning evaluation metrics with application-specific risk characteristics. The benchmark is publicly released with reproducible leaderboards.

probabilistic forecastingproper scoring rulesregression benchmarks

Foundations of the Theory of Performance-Based Ranking

Dec 05, 2024
SP
Sébastien Piérard
🏛️ University of Liège

Existing performance ranking methods in entity evaluation struggle to simultaneously satisfy application-specific preferences and theoretical rigor. Method: This paper establishes the first axiomatic, verifiable general theory framework for performance ranking. Grounded in probability theory and order theory, it formally defines core concepts—including performance objects, satisfaction, and importance—and introduces a performance order satisfying axioms such as ranking consistency, along with constructive procedures for deriving such orders. It further proposes a novel parameterized family of universal ranking scores that unifies classical metrics (e.g., accuracy, recall, F1-score) and rigorously proves that several widely used metrics—including precision—violate the ranking consistency axiom. Contribution/Results: The framework provides the first mathematically rigorous yet practically flexible foundation for performance evaluation in computer vision and machine learning, explicitly characterizing the validity boundaries and intrinsic limitations of reliable ranking metrics.

Establish universal theory for performance-based rankingIntroduce axiomatic definition of performance orderingsPropose parametric family of ranking scores

This study addresses the lack of comparability in peer review scores across research topics at top machine learning conferences, which undermines fairness in paper acceptance decisions. Analyzing 50,289 submissions to ICLR from 2021 to 2026, this work provides the first systematic evidence that papers from different topics with identical review scores exhibit up to an eightfold difference in acceptance probability. The disparity stems from a fundamental flaw in the measurement design of the scoring system—not from individual reviewer bias or cultural differences in scoring practices. Through rigorous statistical modeling and attribution analysis that controls for multiple confounding factors, the authors propose calibrated review signals and topic-conditioned acceptance rates as core metrics for evaluating fairness, offering critical empirical evidence and policy guidance for reforming conference reviewing mechanisms.

acceptance fairnesspeer reviewresearch areas

Latest Papers

What's happening recently
View more

本文评估了不同奖励函数对大语言模型预测性能和行为的影响,比较了五种适当评分规则作为训练目标的效果。

Forecast PerformanceLLM ForecastingModel Calibration

This study addresses the lack of intuitive visualization methods for ordinal regression results, which has hindered their application in fields such as visualization and human-computer interaction. To bridge this gap, the paper proposes, for the first time, the use of modified complementary cumulative distribution function (mCCDF) plots to visualize outputs from cumulative link ordinal regression models. This approach not only fills a critical void in the clear and direct representation of ordinal regression outcomes but also effectively conveys key conclusions consistent with those erroneously derived when treating ordinal variables as continuous. By doing so, the method substantially enhances the interpretability and communicability of model results, offering a principled yet accessible visual framework for practitioners and researchers alike.

CCDF plotsLikert dataordinal regression

This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.

calibrationdistributional validationprobabilistic forecasting

Existing evaluation methods struggle to effectively assess agent performance in subjective, context-dependent, long-horizon enterprise tasks due to their reliance on binary correctness judgments. To address this limitation, this work proposes LH-Bench, a three-pillar evaluation framework that integrates expert-designed rubrics, step-level ground-truth artifact annotations, and pairwise human preference comparisons to enable fine-grained and scalable quantitative assessment. Validation on two real-world scenarios—Figma-to-code translation and procedural content generation—demonstrates that expert-crafted rubrics achieve substantially higher inter-rater agreement than LLM-generated ones (Cohen’s Kappa: 0.60 vs. 0.46), and human preference data significantly align with the framework’s rankings (p < 0.05). The associated dataset has been publicly released.

context-dependent tasksenterprise workflowsLLM assessment

Hot Scholars

JL

Junyang Lin

Qwen Team, Alibaba Group & Peking University
Natural Language ProcessingCross-Modal Representation LearningPretraining
GP

Guansong Pang

Assistant Professor of Computer Science, Singapore Management University
Machine LearningData MiningComputer VisionAnomaly Detection
YC

Yunkang Cao

Hunan University
Visual Anomaly DetectionIndustrial Foundation ModelEmbodied Intelligence
PS

Philip S. Yu

Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy
YC

Yejin Choi

Stanford University / NVIDIA
Natural Language ProcessingDeep LearningArtificial IntelligenceCommonsense Reasoning