Test-Time Compute for Tabular Foundation Models: Mechanisms, Gains, and Limits

๐Ÿ“… 2026-10-08
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates which test-time computation strategies effectively enhance the predictive performance of tabular foundation models. Through a systematic exploration across adaptation, aggregation, and context construction, it proposes DiagScale, a method that fine-tunes minimal parameters via diagonal query-key similarity updates, combined with greedy selective aggregation and attention-guided retrieval. Evaluating models such as TabPFN on the TabArena benchmark, this work reveals the underlying mechanisms by which selective aggregation outperforms uniform averaging. Furthermore, it demonstrates that both parametric adaptation and selective aggregation yield consistent performance gains, with their combination producing superior synergistic effects, albeit at the cost of substantially increased computational overhead.
๐Ÿ“ Abstract
Which forms of test-time compute improve the predictions of strong pretrained tabular foundation models (TFMs)? We systematically study this along three axes: adaptation, aggregation, and context construction. Our evaluation spans modern TFMs across the TabArena benchmark, supplemented by experiments on wide and large-scale tables from OpenML. For adaptation, we introduce DiagScale, a diagonal query-key similarity update. It trains only 0.003-0.03% of model parameters and achieves gains comparable to full fine-tuning across three independently pretrained backbones. For aggregation, both pool composition and selection strategy matter. TabPFN-3 already averages predictions from different preprocessing variants of the same data, and adding more such predictions yields diminishing returns. With a broader pool of 96 configurations, greedy selection reduces error by 2.4% relative to the default predictor, but uniform averaging increases error. For context construction, attention-guided retrieval improves TabPFN-3's predictions on some large tables and supports source pools beyond the full context memory limit. The context expansion methods we test yield no consistent improvement. Taken together, our results suggest that adaptation and selective aggregation yield consistent benchmark-level gains. The benefits of context construction depend more on the task and data regime. Adaptation and aggregation over the same backbone yield further gains when combined, but require substantially more computation than default inference. These trade-offs motivate choosing strategies according to the available computation budget. Code is available at https://github.com/kanghui-learning/test-time-compute-for-tabular-foundation-models.
Problem

Research questions and friction points this paper is trying to address.

Test-Time Compute
Tabular Foundation Models
Adaptation
Aggregation
Context Construction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Compute
Tabular Foundation Models
DiagScale
Greedy Selection
Attention-Guided Retrieval