Score
Designs and implements models and architectures for predicting links when source and target roles are different, including asymmetric graph architectures and separate model towers (e.g., device vs. content); methods include temporally valid message passing, constraining tower inputs for inductive generalization, and retrieval-based graph completion. Also builds evaluation pipelines for link prediction—selecting temporally correct splits and standard metrics, measuring link prediction accuracy and boundary discrimination, and comparing against node-classification and other baselines.
Existing link prediction (LP) evaluation lacks systematic control over critical factors—including network type, geodesic distance distribution, class imbalance, and metric sensitivity—limiting the generalizability of empirical conclusions. Method: We propose the first hypothesis-driven, multidimensional controllable evaluation framework, employing controlled-variable experiments, multi-network benchmarking, and rigorous statistical testing to systematically identify and quantify six previously overlooked sources of evaluation bias. Contribution/Results: We reveal the substantial impact of geodesic distance distribution and class imbalance on the performance of mainstream LP methods; demonstrate the inadequacy of conventional metrics (e.g., AUC) in early-retrieval scenarios; and introduce a hierarchical evaluation paradigm alongside application-oriented best-practice guidelines. This work establishes a methodological foundation for fair, reproducible comparison and reliable deployment of LP methods.
Existing analyses of message-passing graph neural networks (MPNNs) for node and link prediction often rely on unrealistic i.i.d. assumptions, neglecting the critical roles of graph topology, aggregation mechanisms, and loss functions in out-of-distribution inductive generalization. Method: We propose the first unified theoretical framework that systematically models node- and link-level dependencies, supports both inductive and transductive learning, and quantifies—analytically and for the first time—the intrinsic impact of graph topology on generalization error. The framework is agnostic to specific aggregation functions and loss designs. Contribution/Results: Through rigorous theoretical analysis and empirical validation, our framework significantly advances understanding of MPNN generalization behavior, revealing how structural properties govern predictive robustness. It yields interpretable, principled design guidelines for robust graph representation learning, bridging a key gap between theory and practice in geometric deep learning.
This work addresses fully inductive knowledge graph link prediction—requiring zero-shot generalization to both entirely unseen entities and novel relation types at inference time—a setting where existing methods exhibit insufficient cross-dataset generalization. Method: We introduce and formally define the “double permutation-equivariant representation” framework, proving it a necessary condition for this task; we unify the modeling of double equivariance in GNN architectures via group actions and representation theory, integrating theoretical derivation with empirical validation. Contribution/Results: Our analysis reveals a fundamental limitation of double equivariance in cross-domain meta-learning, demonstrating its inadequacy for universal knowledge graph foundation models; we explicitly identify the critical theoretical gap preventing the realization of such cross-domain foundation models—namely, the absence of a principled mechanism for transferring relational abstractions across heterogeneous schema and entity distributions.
In link prediction, conventional GNNs adopt a homogeneous modeling paradigm for all node pairs, ignoring the inherent heterogeneity in their pairwise feature requirements—thereby limiting predictive performance. This work first identifies and quantifies substantial heterogeneity across node pairs in terms of requisite pairwise features (e.g., common neighbors, shortest path distance). To address this, we propose Link-MoE: a dynamic GNN architecture based on Mixture of Experts (MoE), wherein multiple specialized GNNs serve as experts, and a learnable gating mechanism adaptively selects the most suitable expert for each node pair during inference. Evaluated on Pubmed and ogbl-ppa, Link-MoE achieves +18.71% improvement in MRR and +9.59% in Hits@100 over prior state-of-the-art methods. Our core contribution lies in formally characterizing and explicitly modeling pairwise feature heterogeneity—a paradigm shift from the prevailing homogeneous GNN design for link prediction.
Graph Neural Networks (GNNs) exhibit fundamental limitations in modeling structural heuristics—such as Common Neighbors (CN), Adamic-Adar (AA), and Resource Allocation (RA)—for link prediction, primarily due to the inability of set-based neighborhood aggregation to distinguish joint neighborhood structures of node pairs. Method: We propose a GNN framework augmented with trainable node embeddings and conduct systematic evaluation across multiple graph densities, benchmarking against CN/AA/RA heuristics. Contributions/Results: (1) Standard GNNs fail to replicate the predictive performance of classical structural heuristics; (2) Node embeddings implicitly encode link existence information in dense graphs, boosting AUC by up to 8.2%; (3) We establish the first formal linkage between neighborhood aggregation participation and embedding representational capacity. This work establishes theoretical performance bounds for GNN-based link prediction and introduces a novel design paradigm that synergistically integrates structural priors with representation learning.
This study systematically investigates the effectiveness boundaries and practical value of Graph Neural Networks (GNNs) across twelve application domains. Building upon a unified design space, it derives both spectral and spatial formulations of GNNs from first principles, analyzes their expressive power through the lens of the Weisfeiler–Leman test, and evaluates domain-specific graph construction strategies and architectural choices. The work establishes the first cross-domain analytical framework that disentangles genuine performance gains from baseline biases, uncovering common challenges such as heterophily, scaling effects, and deployment gaps. It clarifies the applicability limits of GNNs, highlights the discrepancy between leaderboard-topping models and deployable ones, and offers constraint-aware practical guidelines to address issues including oversmoothing, over-squashing, and distributional shifts.
This paper identifies an information leakage problem in link prediction caused by improper use of the test set during hyperparameter tuning, leading to inflated model performance estimates. To quantify this bias, we propose a novel evaluation metric—Loss Ratio—and conduct a large-scale empirical study across 60 real-world networks using diverse parametric models. Results show that average performance is overestimated by 3.6% on average, with some algorithms exhibiting biases exceeding 15%. Heuristic and random-walk-based methods demonstrate greater robustness to such leakage. The study systematically establishes the necessity of standardized data splitting and evaluation protocols, providing both theoretical grounding and practical guidelines for trustworthy link prediction model assessment.
This work addresses a critical yet overlooked issue in graph neural networks (GNNs): during link prediction training, mini-batch sampling can introduce class composition bias, which—particularly when combined with batch normalization—leads models to learn spurious heuristics. Consequently, the learned representations become misaligned with those beneficial for node classification, challenging the common assumption that GNNs produce task-agnostic, transferable embeddings. The study systematically uncovers this bias for the first time and proposes a correction mechanism that aligns model representations with the intrinsic structural properties of the graph. Experimental results demonstrate that the corrected models focus more on class-relevant features, suggesting that standard training protocols may substantially overestimate the generalization capability of link predictors.
This paper addresses graph representation learning for both static and single-event dynamic networks. Methodologically, it introduces a unified structural-aware embedding framework grounded in latent distance modeling, jointly optimizing homophily, transitivity, and balance within an end-to-end paradigm—thereby eliminating heuristic design and multi-stage pipelines. Notably, it is the first to extend latent distance modeling to single-event dynamic settings, enabling extreme node identification and quantitative assessment of influence dynamics. The key contributions are: (1) a hierarchical, interpretable structural-aware representation; (2) seamless unification of embedding learning across static and dynamic networks; and (3) state-of-the-art performance on community detection, anomaly detection, and temporal influence evaluation—significantly outperforming multi-stage baselines.
Multilayer networks pose significant challenges for embedding learning and link prediction due to their heterogeneous connection types and structural complexity. This work presents a systematic review of existing approaches, introduces a refined taxonomy for modeling methods, and establishes a fair and reproducible evaluation framework. Notably, it designs a novel testing protocol tailored specifically for directed multilayer networks. By doing so, this study establishes the first standardized evaluation paradigm for multilayer network embedding learning, substantially enhancing both link prediction performance and cross-method comparability. The proposed framework advances the field toward more efficient and rigorous research practices.