Distillation of Tabular Foundation Models into Efficient Predictors

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high inference costs of tabular foundation models by proposing an efficient knowledge distillation framework that transfers their capabilities to lightweight student models. Methodologically, it leverages full-label data to construct supervisory signals and expand query coverage, designs a distillation recipe combining complete context with synthetic queries, and integrates neural and tree-based models alongside synthetic data generation techniques. Experimental results demonstrate that the distilled student models achieve Elo rating improvements of 57–98 points, reduce median errors by 4.0%–6.4%, and accelerate inference speed by 3.0× to 21.6×, thereby realizing an effective balance between performance and computational efficiency.
📝 Abstract
Tabular foundation models (TFMs) achieve strong predictive performance through in-context learning, yet repeatedly conditioning on labeled data makes inference expensive. Knowledge distillation can reduce this cost by transferring their predictive ability to lightweight, dataset-specific students. However, the dependence of TFM predictions on both a labeled context and a query introduces two design questions: how to construct teacher supervision and whether expanding query coverage improves distillation. We examine these questions across two TFMs and both neural and tree-based students, and derive an effective distillation recipe. The recipe uses the full labeled training set as teacher context and trains students solely on teacher predictions for observed and synthetic queries. On TabArena, the resulting students outperform their supervised trained tuned-and-ensembled counterparts by 57-98 Elo points. Applied unchanged to TALENT, the same recipe improves matched default students on 236-258 of 300 datasets and reduces median primary error by 4.0-6.4%. The distilled students also achieve median inference speedups of 3.0-21.6 times over their teachers, offering a practical trade-off between predictive performance and repeated inference cost. Code is available at https://github.com/nums-ai/TFM_Distillation .
Problem

Research questions and friction points this paper is trying to address.

Tabular Foundation Models
Knowledge Distillation
Inference Cost
In-context Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Knowledge Distillation
Tabular Foundation Models
In-context Learning
Synthetic Queries
Efficient Inference
🔎 Similar Papers
No similar papers found.