GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic evaluation of large language models (LLMs) on end-to-end graph neural network (GNN) encoding tasks by constructing the first competition-level GNN benchmark. The proposed benchmark encompasses multi-granularity prediction tasks and supports both zero-shot prompting and autonomous agent-based evaluation modes. Through a unified automated pipeline, it enables standardized performance comparisons between humans and LLMs. Experimental results reveal that current LLMs exhibit unstable performance on these tasks and struggle to match top-tier human capabilities. Furthermore, this work releases a dynamic leaderboard and open-sources all associated resources, providing essential infrastructure for advancing research at the intersection of GNNs and LLMs.
📝 Abstract
Large language models (LLMs) have demonstrated strong performance on coding and reasoning benchmarks; however, their ability to solve graph-structured machine learning problems remains largely unexplored. In particular, no benchmark currently evaluates whether LLMs can autonomously solve end-to-end Graph Neural Network (GNN) coding tasks under realistic competition settings. To address this gap, this paper introduces GNN-CB, the first competition-based benchmark for evaluating both humans and LLMs on GNN coding tasks. GNN-CB consists of 18 curated competitions spanning node-, edge-, and graph-level prediction across diverse graph categories, domains, and difficulty tiers. All submissions are evaluated through a unified automated pipeline with hidden test sets and standardized scoring. Human participants solve tasks under controlled competition constraints, while LLMs are evaluated using a frozen zero-shot prompting protocol based on a plan-then-code paradigm with bounded execute-and-repair loops. The benchmark additionally supports both non-agent and autonomous agent-based evaluation within the same protocol. Under our evaluated protocol, LLMs rarely match Human Top performance and show less stable performance across competitions. No single model dominates: a few competitions are won by LLMs, yet humans still hold the top score on most tasks. We release GNN-CB as a living benchmark with automated evaluation infrastructure, dynamic leaderboards, and reproducible execution pipelines. Beyond benchmarking, GNN-CB provides a practice-oriented resource for studying GNN implementation across progressively diverse graph-learning tasks. The benchmark and evaluation framework are publicly available at https://basiralab.github.io/GNN-CB/.
Problem

Research questions and friction points this paper is trying to address.

Graph Neural Networks
Large Language Models
Benchmark
Coding Tasks
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph Neural Network
Competition Benchmark
Large Language Models
Zero-shot Prompting
Autonomous Agent
M
Murad Hossen
University of Houston, Texas, USA
T
Tasneem Selim
Department of Mathematics and Computer Science, Faculty of Science, Alexandria University, Alexandria, Egypt
G
Gurur Gamgam
Bogazici University, Turkey
T
Tuga Yousif
Ankara Yıldırım Beyazıt University, Turkey
A
Abderrahmane Kasmi
ESI, Algiers, Algeria
I
Ikram Aissiou
University of Algiers 1, Algeria
M
Mubaraq Onipede
York St John University London, UK
F
Faran Taimoor Butt
Air University, Islamabad, Pakistan
S
Sanae Zrigui
LSISI, ENSA, Mohammed Premier University, Oujda, Morocco
R
Rosa Y. G. Paccotacya-Yanque
Universidad Católica San Pablo, Arequipa, Peru
I
Ignatius Balayo
Busitema University, Uganda
I
Ikram Elhouiti
University of Laghouat, Algeria
H
Hadil Affes
ISI, University of Tunis El Manar, Tunisia
B
Bijay Adhikari
Tribhuvan University, Nepal
S
Sargam Goyal
IIT Roorkee, India
M
Muhammad Ibrahim Isah
Shobhit Institute of Engineering and Technology, India
M
Mohammad Idrees Bhat
MIT World Peace University, Pune, India
S
Samuel Kangoni Matia
University of Kinshasa, DR Congo
P
Peguy Kem-Meka Tiotsop Kadzue
University of Bertoua, Cameroon; University of the Witwatersrand, Johannesburg, South Africa; African Institute for Mathematical Sciences, Research and Innovation Centre (AIMS RIC), Kigali, Rwanda
M
Maha Trabelsi
ISI, University of Tunis El Manar, Tunisia
Emmanuel Owusu
Emmanuel Owusu
KNUST, Ghana
V
Vinit
NSUT Delhi, India
N
Nour Majdoub
ISIMM Monastir, Tunisia
T
Tamiru Alemnew
Addis Ababa University, Ethiopia
Islem Rekik
Islem Rekik
BASIRA lab
Machine and Deep LearningNeuroimagingNetwork NeurociencePredictive Intelligence in MedicineConnectomics