Thinking in Depth: Retrospective Inference for Tabular Foundation Models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the uneven distribution of prediction refinement across deep Transformer layers in tabular foundation models by proposing Retro. This method explicitly reuses intermediate-layer information through retrospective reasoning and introduces attention residuals alongside query-conditioned gating mechanisms to enable multi-perspective deep refinement. Furthermore, it incorporates adaptive reweighting and element-wise modulation techniques to significantly enhance the utilization efficiency of deep-layer features. Experimental results demonstrate that Retro ranks among the top three on mainstream benchmarks such as TabArena, achieving performance on the Pareto frontier. Consequently, this work establishes an efficient paradigm for deep feature reuse in tabular prediction refinement.
📝 Abstract
Tabular foundation models (TFMs) are pretrained across diverse tabular tasks and make predictions on a new table at inference time using its labeled examples as context. Most recent TFMs perform such in-context prediction with stacked Transformer layers, repeatedly transforming how examples are represented and compared. By tracing individual queries through several strong TFMs, we find that predictive refinement is highly uneven across depth and is often concentrated in later layers. This uneven refinement motivates us to reconsider how intermediate representations are constructed and reused throughout the network. We introduce Retro, a tabular foundation model based on retrospective inference, where later stages can explicitly revisit and recombine intermediate information produced earlier in the network. Retro organizes this process around two complementary operations: which intermediate information to revisit, and how the resulting contextual update should be shaped for each query. Attention Residuals address the former by adaptively reweighting contributions from different depths, while query-conditioned Gated Attention addresses the latter by modulating the attention output element-wise across representation dimensions. Our analysis shows that Retro shifts predictive refinement earlier and more broadly across depth, with different stages revising different subsets of queries in a pattern suggestive of multi-view refinement. Across TabArena, TALENT, and RelArena, Retro ranks among the top three and lies on the Pareto frontier. These results indicate that directly reusing intermediate representations provides a practical way to better exploit depth in TFMs.
Problem

Research questions and friction points this paper is trying to address.

Tabular Foundation Models
In-context Prediction
Predictive Refinement
Intermediate Representations
Transformer Depth
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tabular Foundation Models
Retrospective Inference
Attention Residuals
Query-conditioned Gated Attention
In-context Learning
H
Hao-Run Cai
School of Artificial Intelligence, Nanjing University; National Key Laboratory for Novel Software Technology, Nanjing University
Si-Yang Liu
Si-Yang Liu
Nanjing University
Machine LearningTabular DataLLMs
Z
Zi-Jian Cheng
School of Artificial Intelligence, Nanjing University; National Key Laboratory for Novel Software Technology, Nanjing University
Kun-Yang Yu
Kun-Yang Yu
LAMDA Group, Nanjing University
Machine Learning
J
Jin-Hao Sheng
School of Artificial Intelligence, Nanjing University; National Key Laboratory for Novel Software Technology, Nanjing University
Guo Yu
Guo Yu
University of California, Santa Barbara
High-dimensional statisticsStatistical Machine Learning
C
Chonghan Liu
School of Artificial Intelligence, Nanjing University; National Key Laboratory for Novel Software Technology, Nanjing University
Z
Zhi Zhou
School of Artificial Intelligence, Nanjing University; National Key Laboratory for Novel Software Technology, Nanjing University
Jun-Peng Jiang
Jun-Peng Jiang
Ph.D student at Nanjing University
Tabular Data LearningMultimodal LearningMLLMs
Lan-Zhe Guo
Lan-Zhe Guo
LAMDA Group, Nanjing University
Machine Learning
Han-Jia Ye
Han-Jia Ye
Nanjing University
Machine LearningData MiningMetric LearningMeta-Learning