🤖 AI Summary
This study addresses the failure of fixed-weight interpretability in tabular foundation models caused by label standardization, revealing a joint processing mechanism among contextual labels. Methodologically, overcoming the limitations of conventional derivative analysis confined to standardized ranges, this work introduces a novel mathematical certificate relying solely on standardized label predictions, validated through attention evaluation and experiments across five public models. The results demonstrate that nonlinear joint processing—where altering a single label influences others—is universally present across all tested models. Furthermore, this property emerges during training and is primarily mediated by the attention mechanism. These findings offer a new perspective for interpretability research in tabular foundation models.
📝 Abstract
Linear regression and kernel smoothing offer tractable explanations of in-context learning: in both, the features determine the weight assigned to each context label. However, whether this fixed-weight account describes pretrained tabular foundation models (TFMs) remains unclear. Testing this account using derivatives runs into a standardization trap: public TFM packages standardize the labels before the model sees them, yet ordinary derivatives also reflect behavior outside the set of standardized labels, making a model appear nonlinear even when every prediction it makes agrees with a fixed-weight map. We propose two certificates that depend only on predictions at standardized labels and can reject two distinct explanations: fixed-weight prediction and sums of independent nonlinear label transformations. Across the five public TFMs that we evaluate, our certificates show that changing one context label alters how other labels influence the prediction, a behavior we call joint processing. We further find that joint processing emerges with training and that attention scores carry most of the measured interaction. Together, these findings motivate TFM explanations that account for how context labels change the influence of individual examples.