When Does Synthetic Relational Data Teach Models to Use Relations? Tracing Predictive Structure from Pretraining Data to Model Behavior

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear relationship between properties of synthetic data and model learning of relational structures, which currently disconnects performance improvements from their underlying mechanisms. To bridge this gap, we construct a causal chain spanning data properties, learning mechanisms, and downstream behaviors. By integrating Relational Transformers, data attribution analysis, and foreign key intervention experiments, we systematically trace the mapping pathway from pretraining data to model behavior. Our findings reveal that the necessity for cross-table prediction is the core factor driving relational computation. Furthermore, we demonstrate that the RelDiff generator establishes its advantage by inducing unique serial cross-table pathways; disrupting this mechanism entirely eliminates its performance gains. These insights provide explicit theoretical guidance for the principled design of synthetic relational data.
📝 Abstract
Relational foundation models are increasingly pretrained on synthetic databases, yet downstream benchmarks reveal little about why one synthetic corpus produces a better model than another. In particular, strong performance may arise from realistic row-level statistics without the model ever learning to use relational structure. We study this as a data-attribution problem: which property of synthetic pretraining data induces relational computation? Using four Relational Transformer checkpoints trained with the same architecture, initialization, objective, and compute budget on corpora produced by four relational data generators, we trace a measurable property of the data to learned computation and downstream behavior. We hypothesize that relational mechanisms emerge when cross-table information is predictively necessary for the masked-cell pretraining objective. RelDiff exhibits by far the largest predictive gain from foreign-key-linked parents, and its corresponding model is uniquely sensitive to foreign-key interventions on unseen databases. This dependence survives a random-initialization control, grows monotonically with the fraction of corrupted links, and localizes to a serial cross-table pathway. Finally, disrupting the same mechanism during downstream inference removes RelDiff's advantage on relational tasks while leaving structure-insensitive models nearly unchanged. These results connect a property of synthetic training data to a learned mechanism and, through intervention, to downstream behavior.
Problem

Research questions and friction points this paper is trying to address.

synthetic relational data
relational foundation models
data attribution
relational computation
pretraining
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data Attribution
Relational Foundation Models
Synthetic Pretraining Data
Mechanistic Interpretability
Causal Intervention
💼 Related Jobs
No related jobs found.