TICDA: Tabular In-Context Data Attribution

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of efficiently quantifying the influence of demonstration examples on predictions during in-context learning with tabular foundation models, where conventional attribution methods encounter significant computational bottlenecks. To this end, this work proposes TICDA, a framework that directly measures the data attribution contribution of demonstrations to prediction outcomes by constructing linear surrogate models and conducting latent space embedding analysis. Notably, this approach requires only a single forward pass, eliminating the need for parameter updates or multiple inference iterations. The proposed TICDA framework achieves low-cost, high-precision data influence quantification while effectively balancing attribution accuracy with inference efficiency. Extensive evaluations demonstrate its superior performance across diverse downstream tasks, including error detection, context selection, and active learning, establishing it as a practical solution for interpretable in-context learning in tabular domains.
📝 Abstract
Tabular foundation models (TFMs) achieve strong predictive performance by conditioning on labeled demonstrations provided in context, without any parameter update. Yet how individual demonstrations shape a given prediction remains poorly understood. This gap matters in practice: the context is often assembled from whatever labeled data is available, potentially leading to the inclusion of mislabeled, redundant, or low-quality examples that degrade performance. Standard data attribution methods do not transfer to the TFM setting: resampling-based approaches such as DemoShapley require a combinatorial number of forward passes, and gradient-based estimators such as influence functions require computing training point's effect on the model parameters, which in-context learning never updates. We introduce TICDA, a method that measures the influence of every demonstration in the context directly from linear surrogates trained on TFM latent embeddings, in a single forward pass and at negligible cost. We show that TICDA offers the best compromise against competitors across four tasks: detecting labeling errors, curating context to preserve predictive accuracy while lowering inference cost, producing attribution scores that transfer across TFMs, and supporting an acquisition strategy for efficient active learning.
Problem

Research questions and friction points this paper is trying to address.

Tabular Foundation Models
In-Context Learning
Data Attribution
Demonstration Influence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tabular Foundation Models
In-Context Learning
Data Attribution
Linear Surrogates
Latent Embeddings
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yacine Benihaddadene
Ekimetrics
M
Milan Bhan
Ekimetrics
E
Eliot Dugelay
ETH Zurich
M
Mohammed Jawhar
Ekimetrics
B
Benjamin Wong
Ekimetrics
Nicolas Chesneau
Nicolas Chesneau
Unknown affiliation
D
Duong Nguyen