Architecture Alignment With Sparse Priors in Tabular Foundation Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear mechanisms of irrelevant feature suppression in tabular foundation models arising from architectural discrepancies. To this end, it proposes an "architecture-prior alignment" theory, revealing that addressable feature axes provide inductive biases for task-adaptive relevance reasoning. Employing controlled Transformer training, Bayesian analysis, and attention intervention techniques, this work systematically compares row-token and alternating-axis architectures under sparsity priors. Results demonstrate that the alternating-axis architecture significantly outperforms the row-token counterpart on sparse tasks, yielding predictions that more closely approximate the Bayesian optimal solution. These findings validate the critical role of architectural alignment in effective feature gating.
📝 Abstract
Tabular foundation models (TFMs) are increasingly popular because they deliver strong predictions on new datasets through in-context learning, without task-specific training or extensive tuning. Yet released TFMs differ simultaneously in their pretraining priors, architectures, and objectives, obscuring their respective inductive biases. We therefore examine one concrete capability: irrelevant-feature suppression. Across synthetic tasks and real-world datasets, adding null features causes substantially greater predictive degradation in the row-token model TabDPT, whereas the cell-token alternating-axis model TabPFN v2 and other TFMs remain comparatively stable. This gap motivates us to ask whether architecture contributes to irrelevant-feature suppression. Because released TFMs remain confounded by other design choices, we train streamlined row-token and alternating-axis transformers under identical sparse-to-dense linear priors. Exact Bayes analysis shows that sparse prediction requires context-dependent feature gating, whereas the dense endpoint requires only uniform feature weighting. Consistent with this distinction, the alternating-axis model is substantially closer to the Bayesian optimal predictor on sparse tasks, while the architecture gap becomes negligible on dense tasks; almost all of the sparse gap arises from linear coefficient-estimation error. Finally, in both the controlled model and frozen TabPFN v2, we examine the effect of interventions on the feature-attention outputs on the linear coefficients, finding evidence of task-dependent selective routing of computation through feature-indexed pathways. Together, these results support architecture-prior alignment: preserving an addressable feature axis provides an inductive bias for task-adaptive relevance inference. Code is available at https://github.com/Tianqi-Zhao/ArchitecturePriorTFMs.
Problem

Research questions and friction points this paper is trying to address.

Tabular Foundation Models
Irrelevant-feature Suppression
Architecture Alignment
Sparse Priors
Inductive Bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tabular Foundation Models
Architecture-Prior Alignment
Sparse Priors
Feature Gating
Alternating-Axis Transformer
🔎 Similar Papers
No similar papers found.