From Behavior to Provenance: Attributing Tabular Foundation Models to Synthetic Pretraining Data

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck in verifying training data attribution for tabular foundation models, where causal influences remain unobservable. We propose controlled synthetic pretraining as a novel paradigm for attribution experimentation. By leveraging O'PRIOR to generate synthetic pretraining tasks enriched with provenance information, and integrating behavior-conditioned attribution with counterfactual retraining, we construct a verifiable attribution testbed that enables dual validation of mechanism-level consistency and task-level faithfulness. Experimental results demonstrate that removing the top 5% of highly attributed tasks yields a 0.013 decrease in ROC-AUC, significantly outperforming random baselines. These findings confirm the effectiveness of synthetic provenance in achieving verifiable contribution attribution for tabular foundation models.
📝 Abstract
Training-data attribution aims to identify which training examples shape model behavior, yet validating such claims is difficult because causal training influence is rarely observable. We argue that controlled synthetic pretraining makes attribution experimentally testable. Using O'PRIOR, a provenance-rich synthetic task generator for tabular foundation models, we construct a testbed in which every pretraining task carries explicit lineage over structural mechanisms, missingness, confounding, shortcuts, and distribution shift. We combine behavior-conditioned attribution with counterfactual retraining and provenance-aware interventions to test both task-level faithfulness and mechanism-level consistency. On held-out real tasks, removing the top-attributed 5% of synthetic tasks decreases mean ROC-AUC by 0.013, compared with 0.002$\pm$0.004 under random removal, while removing bottom-attributed tasks improves performance by 0.003. Within shortcut-provenance tasks, targeted removal yields an effect of 0.043 versus 0.016 for matched random removal. Provenance discrimination is more modest by ranking AUROC (0.55-0.62), despite substantial top-k enrichment, revealing that provenance association and interventional faithfulness need not coincide. Our results establish synthetic provenance as a controlled setting for verifiable contributive attribution
Problem

Research questions and friction points this paper is trying to address.

Training-data attribution
Tabular foundation models
Synthetic pretraining data
Data provenance
Model behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-data attribution
Synthetic pretraining
Tabular foundation models
Data provenance
Counterfactual retraining
🔎 Similar Papers
2024-06-16International Conference on Learning RepresentationsCitations: 12
💼 Related Jobs
No related jobs found.