🤖 AI Summary
This study addresses the bottleneck in verifying training data attribution for tabular foundation models, where causal influences remain unobservable. We propose controlled synthetic pretraining as a novel paradigm for attribution experimentation. By leveraging O'PRIOR to generate synthetic pretraining tasks enriched with provenance information, and integrating behavior-conditioned attribution with counterfactual retraining, we construct a verifiable attribution testbed that enables dual validation of mechanism-level consistency and task-level faithfulness. Experimental results demonstrate that removing the top 5% of highly attributed tasks yields a 0.013 decrease in ROC-AUC, significantly outperforming random baselines. These findings confirm the effectiveness of synthetic provenance in achieving verifiable contribution attribution for tabular foundation models.
📝 Abstract
Training-data attribution aims to identify which training examples shape model behavior, yet validating such claims is difficult because causal training influence is rarely observable. We argue that controlled synthetic pretraining makes attribution experimentally testable. Using O'PRIOR, a provenance-rich synthetic task generator for tabular foundation models, we construct a testbed in which every pretraining task carries explicit lineage over structural mechanisms, missingness, confounding, shortcuts, and distribution shift. We combine behavior-conditioned attribution with counterfactual retraining and provenance-aware interventions to test both task-level faithfulness and mechanism-level consistency. On held-out real tasks, removing the top-attributed 5% of synthetic tasks decreases mean ROC-AUC by 0.013, compared with 0.002$\pm$0.004 under random removal, while removing bottom-attributed tasks improves performance by 0.003. Within shortcut-provenance tasks, targeted removal yields an effect of 0.043 versus 0.016 for matched random removal. Provenance discrimination is more modest by ranking AUROC (0.55-0.62), despite substantial top-k enrichment, revealing that provenance association and interventional faithfulness need not coincide. Our results establish synthetic provenance as a controlled setting for verifiable contributive attribution