🤖 AI Summary
This study addresses the uneven distribution of prediction refinement across deep Transformer layers in tabular foundation models by proposing Retro. This method explicitly reuses intermediate-layer information through retrospective reasoning and introduces attention residuals alongside query-conditioned gating mechanisms to enable multi-perspective deep refinement. Furthermore, it incorporates adaptive reweighting and element-wise modulation techniques to significantly enhance the utilization efficiency of deep-layer features. Experimental results demonstrate that Retro ranks among the top three on mainstream benchmarks such as TabArena, achieving performance on the Pareto frontier. Consequently, this work establishes an efficient paradigm for deep feature reuse in tabular prediction refinement.
📝 Abstract
Tabular foundation models (TFMs) are pretrained across diverse tabular tasks and make predictions on a new table at inference time using its labeled examples as context. Most recent TFMs perform such in-context prediction with stacked Transformer layers, repeatedly transforming how examples are represented and compared. By tracing individual queries through several strong TFMs, we find that predictive refinement is highly uneven across depth and is often concentrated in later layers. This uneven refinement motivates us to reconsider how intermediate representations are constructed and reused throughout the network. We introduce Retro, a tabular foundation model based on retrospective inference, where later stages can explicitly revisit and recombine intermediate information produced earlier in the network. Retro organizes this process around two complementary operations: which intermediate information to revisit, and how the resulting contextual update should be shaped for each query. Attention Residuals address the former by adaptively reweighting contributions from different depths, while query-conditioned Gated Attention addresses the latter by modulating the attention output element-wise across representation dimensions. Our analysis shows that Retro shifts predictive refinement earlier and more broadly across depth, with different stages revising different subsets of queries in a pattern suggestive of multi-view refinement. Across TabArena, TALENT, and RelArena, Retro ranks among the top three and lies on the Pareto frontier. These results indicate that directly reusing intermediate representations provides a practical way to better exploit depth in TFMs.