GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation
This study addresses the lack of systematic evaluation of large language models (LLMs) on end-to-end graph neural network (GNN) encoding tasks by constructing the first competition-level GNN benchmark. The proposed benchmark encompasses multi-granularity prediction tasks and supports both zero-shot prompting and autonomous agent-based evaluation modes. Through a unified automated pipeline, it enables standardized performance comparisons between humans and LLMs. Experimental results reveal that current LLMs exhibit unstable performance on these tasks and struggle to match top-tier human capabilities. Furthermore, this work releases a dynamic leaderboard and open-sources all associated resources, providing essential infrastructure for advancing research at the intersection of GNNs and LLMs.