🤖 AI Summary
Existing heterogeneous graph neural networks (HGNNs) heavily rely on node/edge type labels for parameterization, leading to poor semantic generalizability, limited cross-type knowledge transfer, and weak interpretability. To address this, we propose the first integration of a Mixture-of-Experts (MoE) mechanism into the Heterogeneous Graph Transformer (HGT), introducing a type-agnostic, semantics-driven expert routing scheme. Specifically, we randomly mask type embeddings during training to attenuate reliance on superficial type labels, enabling experts to specialize according to intrinsic semantic patterns rather than predefined types. Evaluated on link prediction across IMDB, ACM, and DBLP, our method significantly outperforms standard HGT and type-aware MoE baselines. It achieves superior generalizability, higher computational efficiency, and enhanced interpretability—offering a novel lightweight, semantics-adaptive architectural paradigm for heterogeneous graph modeling.
📝 Abstract
A common practice in heterogeneous graph neural networks (HGNNs) is to condition parameters on node/edge types, assuming types reflect semantic roles. However, this can cause overreliance on surface-level labels and impede cross-type knowledge transfer. We explore integrating Mixture-of-Experts (MoE) into HGNNs--a direction underexplored despite MoE's success in homogeneous settings. Crucially, we question the need for type-specific experts. We propose Homogeneous Expert Routing (HER), an MoE layer for Heterogeneous Graph Transformers (HGT) that stochastically masks type embeddings during routing to encourage type-agnostic specialization. Evaluated on IMDB, ACM, and DBLP for link prediction, HER consistently outperforms standard HGT and a type-separated MoE baseline. Analysis on IMDB shows HER experts specialize by semantic patterns (e.g., movie genres) rather than node types, confirming routing is driven by latent semantics. Our work demonstrates that regularizing type dependence in expert routing yields more generalizable, efficient, and interpretable representations--a new design principle for heterogeneous graph learning.