🤖 AI Summary
This study investigates whether Graph Foundation Models (GFMs) can replace traditional supervised learning while balancing predictive quality and computational cost. To this end, we construct a unified evaluation benchmark that systematically compares six GFMs against fifteen supervised methods across 51 datasets. We introduce an end-to-end workflow measurement encompassing adaptation, training, and inference stages, ensuring fair comparisons through shared data splits, validation-based model selection, controlled hyperparameter search, and multi-metric evaluation. Our empirical findings demonstrate that fine-tuned Graph Neural Networks generally outperform other approaches, whereas directly reusing pretrained models without adaptation does not yet offer a reliable pathway to performance or efficiency gains. These insights provide critical guidance for the practical deployment of GFMs.
📝 Abstract
Can a pretrained graph model replace training and tuning a separate predictor for each dataset? Answering this requires evaluating prediction quality alongside computational cost. We present NodeGround, a node classification benchmark that puts graph foundation models (GFMs) and dataset-specific supervised learning under a common evaluation framework. The benchmark spans 51 datasets and evaluates six GFMs alongside 15 supervised methods under two label-availability regimes. Shared data partitions, validation-only model selection, controlled hyperparameter searches, and multiple predictive metrics make comparisons systematic, while workflow measurements account for adaptation, training, tuning, and inference. The results favor carefully tuned graph neural networks overall. GraphPFN reaches third place by Elo when more labels are available, yet its relative strengths vary substantially with dataset properties. Efficiency comparisons further qualify the benefits of pretrained reuse: GVT and GraphPFN appear on the Pareto frontiers when supervised methods are represented by their default and fully tuned configurations. Adding intermediate tuning budgets removes this advantage for GVT and leaves GraphPFN extending the estimated frontier in the label-rich setting alone. Thus, reusing pretrained parameters does not yet provide a broadly reliable route to either stronger predictions or cheaper workflows. We release the evaluation pipeline, run-level records, and an open leaderboard at https://github.com/nums-ai/nodeground.