🤖 AI Summary
Modeling tabular data often neglects external dependencies among samples—e.g., sparse relational graphs—limiting the capture of non-local, structured inter-sample relationships beyond local neighborhood constraints imposed by graph neural networks (GNNs).
Method: We propose the Relation-Aware Transformer (RA-Transformer), which explicitly encodes prior pairwise relations via a learnable relation-aware attention correction term. It integrates relation-graph regularization and synthetic data augmentation to jointly leverage structural priors and data diversity.
Contribution/Results: We systematically evaluate RA-Transformer on regression and causal inference tasks—specifically treatment effect estimation. Experiments demonstrate consistent and significant improvements over strong baselines (e.g., gradient-boosted trees) across synthetic benchmarks, real-world datasets, and the IHDP benchmark. The results validate its effectiveness in modeling sparse, non-local dependencies, establishing a novel graph-augmented paradigm for tabular data modeling.
📝 Abstract
Deep learning models for tabular data typically do not allow for imposing a graph of external dependencies between samples, which can be useful for accounting for relatedness in tasks such as treatment effect estimation. Graph neural networks only consider adjacent nodes, making them difficult to apply to sparse graphs. This paper proposes several solutions based on a modified attention mechanism, which accounts for possible relationships between data points by adding a term to the attention matrix. Our models are compared with each other and the gradient boosting decision trees in a regression task on synthetic and real-world datasets, as well as in a treatment effect estimation task on the IHDP dataset.