🤖 AI Summary
This study addresses the limited generalizability of existing graph foundation models across modalities, feature spaces, and tasks by proposing Wander, a unified framework that leverages random walks as a universal interface. By integrating a joint pretraining strategy with Bayesian optimal approximation theory, Wander enables general-purpose pretraining on heterogeneous graphs and structural context expansion during inference. Furthermore, it introduces a novel graph completion paradigm from a prior-prediction perspective, facilitating positive transfer and capability composition across diverse graph tasks using a single checkpoint. Extensive experiments demonstrate that Wander achieves state-of-the-art performance on node classification, link prediction, and other benchmarks, thoroughly validating its versatility and effectiveness.
📝 Abstract
Graph foundation models aim to transfer across graphs, feature spaces, relational schemas, and prediction tasks, yet existing approaches typically generalize only within particular graph modalities or tasks. We propose Wander, a graph foundation model designed to operate across these settings within a single pretrained checkpoint. Following the prior-predictive perspective, we formulate graph learning as completion of a partially observed graph. We realize this task-general view through a common interface based on random walks, allowing the same model to operate across homogeneous and multi-relational graphs with varying features, labels, and relational schemas. Wander can increase its structural context at inference time without changing its learned parameters and, under suitable assumptions, universally approximates the corresponding Bayes-optimal predictor on bounded connected graphs. Empirically, a single pretrained checkpoint achieves state-of-the-art or highly competitive results across node classification, homogeneous link prediction, and knowledge-graph link prediction. Moreover, joint pretraining across graph modalities and tasks preserves performance in specialized settings while enabling positive transfer and the composition of separately learned capabilities at inference time.