Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models

📅 2026-07-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates the out-of-distribution (OOD) robustness of nine tabular foundation models under real-world distribution shifts driven by label, socioeconomic, and geographic factors. Leveraging the HELOC, Voting, and Childhood Lead datasets from the TableShift benchmark, it presents the first empirical analysis of OOD performance across diverse shifts for models such as TabPFN, TabICL, Mitra, LimiX, and TabFM. The findings reveal consistent performance degradation ranging from 0.003 to 0.060 across all models, with a strong positive correlation between in-distribution and OOD performance. Moreover, the work highlights that high-performing models often incur prohibitive computational and memory costs, limiting their practical deployability. By extending the OOD evaluation framework for tabular data, this research uncovers a fundamental tension between model scalability and real-world resource constraints.
📝 Abstract
Tabular Foundation Models (TFMs) have emerged as novel approaches for tabular predictive tasks, demonstrating competitive predictive performance to ensemble tree-based models. Most TFMs are trained and evaluated on independent and identically distributed data, but this assumption changes in real-world scenarios due to distribution shifts, which compromise the robustness of models. Limited research has been conducted of TFMs under distribution shifts. We present an empirical evaluation of Out-Of-Distribution (OOD) performance of nine TFMs, spanning diverse pre-training strategies and architectures: TabPFNv2, TabPFNv2.5, TabPFNv2.6, TabPFNv3, TabICL, TabICLv2, Mitra, LimiX and TabFM. Three real-world datasets from the TableShift study were considered (HELOC, Voting, Childhood Lead), covering label, socioeconomic, and geographic shift types. Our results show that all evaluated TFMs degrade systematically under distribution shift regardless of pre-training strategy, with shift gaps ranging from 0.003 to 0.060 depending on shift type. The relationship between in-distribution and OOD predictive performance documented for classical tabular models extends into TFMs. We also identified a scalability gap, as high-performing models demand significant memory and computational resources beyond what standard deployment infrastructure can support. This study extends existing benchmarks for OOD in tabular data, providing evidence to support their adoption in high-stakes domains characterized by structural distribution shifts.
Problem

Research questions and friction points this paper is trying to address.

Tabular Foundation Models
Out-Of-Distribution
distribution shift
robustness
empirical evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tabular Foundation Models
Out-Of-Distribution
Distribution Shift
Empirical Evaluation
Scalability Gap
M
Malena Loza
Colegio de Ciencias e Ingenierías, Universidad San Francisco de Quito (USFQ), Quito, Ecuador
D
David Chushig-Muzo
Department of Signal Theory and Communications, Telematics and Computing Systems, Rey Juan Carlos University, Madrid, Spain
E
Eva Milara
Department of Signal Theory and Communications, Telematics and Computing Systems, Rey Juan Carlos University, Madrid, Spain
L
Luis Bote-Curiel
Department of Signal Theory and Communications, Telematics and Computing Systems, Rey Juan Carlos University, Madrid, Spain
L
Luis Estrada-Petrocelli
Facultad de Ingeniería, Universidad Latina de Panamá, Ciudad de Panamá, Panamá
Felipe Grijalva
Felipe Grijalva
Associate Professor at USFQ
Signal ProcessingSpatial AudioMachine LearningComputer VisionAssistive Technologies