🤖 AI Summary
This study addresses the inefficiency of dataset-specific tuning and AutoML search in tabular data prediction by proposing a 400-million-parameter tabular foundation model. The model reformulates prediction as in-context learning, enabling zero-shot calibrated predictions through a single forward pass. Trained entirely on synthetic causal models, it generalizes to real-world scenarios without task-specific fine-tuning. Furthermore, the approach incorporates large language model-guided data processing, multi-view feature integration, and post-hoc calibration techniques. Evaluated across 51 benchmarks, the proposed model consistently outperforms fully tuned AutoML pipelines, achieving state-of-the-art performance and ranking first overall.
📝 Abstract
Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We present TabFM, a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning. TabFM produces calibrated zero-shot predictions in a single forward pass without task-specific tuning. Trained entirely on synthetic tables generated from structural causal models, TabFM learns general tabular representations that transfer zero-shot to real-world tasks. Across all 51 benchmark datasets in TabArena (38 classification and 13 regression), zero-shot TabFM ranks first among default tabular foundation models and outperforms tuned AutoML pipelines. Two extensions over the same frozen weights improve performance further on both tracks: multi-view feature expansion with ensembling and post-hoc calibration (TabFM+), and LLM-guided, dataset-specific data processing and feature engineering (TabFM-Auto).