🤖 AI Summary
This study addresses the issue in relational deep learning where two temporal signals—record age and observation interval—are often conflated or processed in isolation. To this end, we propose a two-stage self-supervised pretraining framework that innovatively jointly encodes dynamic record ages and fixed observation intervals by integrating multi-scale and rotary time embeddings. Built upon Graph Neural Network and Transformer backbones, the method introduces three self-supervised objectives, including historical relation recovery, to explicitly model heterogeneous temporal dynamics within heterogeneous graphs. Extensive evaluations on the RelBench benchmark demonstrate that our approach achieves a 3.24% improvement over non-pretrained baselines and outperforms supervised training under identical configurations by 1.06%–3.02%, significantly enhancing downstream task performance.
📝 Abstract
Relational deep learning models database rows and foreign-key links as a heterogeneous graph for prediction from record attributes and relational context. These graphs contain two distinct temporal signals: record age changes with the prediction cutoff, while intervals between observed records remain fixed. Prior work often treats time as a single signal or studies temporal representation and pretraining separately. We investigate how explicitly encoding both signals affects temporal pretraining for downstream tasks. Our framework combines Multi-scale Time Encoding, which captures record age using learnable time scales and type-specific projections, with Rotary Time Encoding, which represents signed inter-record intervals through rotary transformations during graph propagation. We pair these encodings with three self-supervised objectives: historical relation recovery, horizon-aware future relation activity prediction, and temporal subgraph contrast. All inputs respect their observation cutoffs. Pretraining proceeds in two stages: subgraph contrast first learns neighborhood representations, followed by refinement through either relation recovery or future activity prediction. We evaluate on five RelBench datasets across 11 classification and regression tasks using heterogeneous GNN and graph Transformer backbones. With both encodings, the best evaluated staged schedules improve over supervised training with the same encodings by 3.02% and 1.06% on the two backbones, respectively, and over controls without pretraining or either encoding by 3.24% and 2.37%.