ScoringBench: A Benchmark for Evaluating Tabular Foundation Models with Proper Scoring Rules

📅 2026-03-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing tabular foundation models are predominantly evaluated using point-estimate metrics such as RMSE and R², which fail to capture tail behavior of predictive distributions and cannot accommodate the asymmetric risk modeling required in high-stakes domains like finance and clinical decision-making. To address this gap, this work proposes ScoringBench—an open-source benchmark that systematically integrates a diverse set of proper scoring rules, including CRPS, CRLS, interval score, and energy score, enabling comprehensive assessment of probabilistic prediction quality alongside traditional metrics. Empirical results demonstrate that model rankings vary substantially depending on the chosen scoring rule, and no single pretraining objective consistently dominates across all criteria, underscoring the necessity of aligning evaluation metrics with application-specific risk characteristics. The benchmark is publicly released with reproducible leaderboards.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationReasoning under Uncertainty: Relational Probabilistic ModelsNatural Language Processing: Safety and Robustness

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsUser Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating success
📝 Abstract
Tabular foundation models such as TabPFN and TabICL already produce full predictive distributions yet prevailing regression benchmarks evaluate them almost exclusively via point estimate metrics RMSE R2 These aggregate measures often obscure model performance in the tails of the distribution a critical deficit for high stakes decision making in domains like finance and clinical research where asymmetric risk profiles are the norm We introduce ScoringBench an open benchmark that computes a comprehensive suite of proper scoring rules like CRPS CRLS Interval Score Energy Score weighted CRPS and Brier Score alongside standard point metrics providing a richer picture of probabilistic forecast quality We evaluate realTabPFNv2.5 fine tuned with different scoring rule objectives and TabICL relative to untuned realTabPFNv2.5 across a suite of regression benchmarks Our results confirm that model rankings depend on the chosen scoring rule and that no single pretraining objective is universally optimal This demonstrates that for applications sensitive to extreme events the choice of evaluation metric is as much a domain specific requirement as the data itself ScoringBench is available at https://github.com/jonaslandsgesell/ScoringBench A live preview of the current leaderboard is available at https://scoringbench.bolt.host The leaderboard is maintained via git pull requests to ensure transparency traceability agility and reproducibility
Problem

Research questions and friction points this paper is trying to address.

tabular foundation models
proper scoring rules
probabilistic forecasting
regression benchmarks
tail performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

proper scoring rules
tabular foundation models
probabilistic forecasting
ScoringBench
distributional evaluation
💼 Related Jobs
No related jobs found.
J
Jonas Landsgesell
P
Pascal Knoll