🤖 AI Summary
This study systematically evaluates the intrinsic validity of heterogeneous graph neural networks (HGNNs), addressing prevalent implicit assumptions and the lack of causal validation in the field. We propose the first causal effect estimation framework for HGNNs, integrating counterfactual analysis, minimal sufficient adjustment set identification, cross-method consistency checks, and sensitivity analysis. Conducting large-scale replication experiments across 21 datasets and 20 baseline models, we find that heterogeneous information exerts a statistically significant positive causal effect on model performance—primarily by enhancing node representation homogeneity and mitigating distributional shift, thereby improving classification discriminability; in contrast, model complexity exhibits no significant causal contribution. The implementation is publicly available, establishing a causal benchmark for interpretable evaluation and architecture design of HGNNs.
📝 Abstract
Graph neural networks (GNNs) have achieved remarkable success in node classification. Building on this progress, heterogeneous graph neural networks (HGNNs) integrate relation types and node and edge semantics to leverage heterogeneous information. Causal analysis for HGNNs is advancing rapidly, aiming to separate genuine causal effects from spurious correlations. However, whether HGNNs are intrinsically effective remains underexamined, and most studies implicitly assume rather than establish this effectiveness. In this work, we examine HGNNs from two perspectives: model architecture and heterogeneous information. We conduct a systematic reproduction across 21 datasets and 20 baselines, complemented by comprehensive hyperparameter retuning. To further disentangle the source of performance gains, we develop a causal effect estimation framework that constructs and evaluates candidate factors under standard assumptions through factual and counterfactual analyses, with robustness validated via minimal sufficient adjustment sets, cross-method consistency checks, and sensitivity analyses. Our results lead to two conclusions. First, model architecture and complexity have no causal effect on performance. Second, heterogeneous information exerts a positive causal effect by increasing homophily and local-global distribution discrepancy, which makes node classes more distinguishable. The implementation is publicly available at https://github.com/YXNTU/CausalHGNN.