๐ค AI Summary
Graph Neural Networks (GNNs) exhibit fundamental limitations in modeling structural heuristicsโsuch as Common Neighbors (CN), Adamic-Adar (AA), and Resource Allocation (RA)โfor link prediction, primarily due to the inability of set-based neighborhood aggregation to distinguish joint neighborhood structures of node pairs.
Method: We propose a GNN framework augmented with trainable node embeddings and conduct systematic evaluation across multiple graph densities, benchmarking against CN/AA/RA heuristics.
Contributions/Results: (1) Standard GNNs fail to replicate the predictive performance of classical structural heuristics; (2) Node embeddings implicitly encode link existence information in dense graphs, boosting AUC by up to 8.2%; (3) We establish the first formal linkage between neighborhood aggregation participation and embedding representational capacity. This work establishes theoretical performance bounds for GNN-based link prediction and introduces a novel design paradigm that synergistically integrates structural priors with representation learning.
๐ Abstract
This paper explores the ability of Graph Neural Networks (GNNs) in learning various forms of information for link prediction, alongside a brief review of existing link prediction methods. Our analysis reveals that GNNs cannot effectively learn structural information related to the number of common neighbors between two nodes, primarily due to the nature of set-based pooling of the neighborhood aggregation scheme. Also, our extensive experiments indicate that trainable node embeddings can improve the performance of GNN-based link prediction models. Importantly, we observe that the denser the graph, the greater such the improvement. We attribute this to the characteristics of node embeddings, where the link state of each link sample could be encoded into the embeddings of nodes that are involved in the neighborhood aggregation of the two nodes in that link sample. In denser graphs, every node could have more opportunities to attend the neighborhood aggregation of other nodes and encode states of more link samples to its embedding, thus learning better node embeddings for link prediction. Lastly, we demonstrate that the insights gained from our research carry important implications in identifying the limitations of existing link prediction methods, which could guide the future development of more robust algorithms.