Score
Designs and implements graph-attention network components that operate on two parallel feature streams—such as dense geometric and semantic representations—by performing edge-guided, relational message passing and attention across nodes and edges. Builds adaptive gated and contextual fusion mechanisms that combine complementary modality semantics while avoiding feature interference during aggregation.
This study addresses the lack of a systematic survey on the integration of attention mechanisms with graph neural networks (GNNs), which has hindered a clear understanding of their developmental trajectory. To bridge this gap, the work proposes a novel two-level taxonomic framework that organizes the field both historically and architecturally. At the upper level, it delineates three chronological phases: Graph Recurrent Attention Networks, Graph Attention Networks, and Graph Transformers. The lower level systematically catalogs representative models within each phase and compares their key characteristics. Through comprehensive literature review and taxonomy-based analysis, the paper elucidates the evolutionary pathway of attention in GNNs, clarifies the strengths and limitations of existing approaches, identifies open challenges, and outlines promising future directions. An accompanying open-source repository is provided to foster ongoing community research.
This work systematically investigates the performance-efficiency trade-offs between dense and sparse attention mechanisms in graph transformers for modeling long-range node dependencies. We conduct controlled experiments across diverse graph structures—including social networks, molecular graphs, and citation networks—to quantitatively characterize their expressive power, computational cost, memory footprint, and scalability. Guided by empirical findings, we propose a “structure-aware sparsification” design principle: hybrid local-global sparse attention for high-clustering-coefficient graphs, and selective retention of global dense connections for small-diameter, homogeneous graphs. Experiments demonstrate that our strategy preserves over 95% of dense-attention accuracy while reducing inference latency by 3.2× and GPU memory consumption by 68%. The study reveals a strong coupling between graph topological properties and attention paradigm efficacy, establishing an evidence-based foundation for architecture search and hardware-aware optimization of efficient graph transformers.
Standard graph attention networks struggle with unreliable node features and fixed sharpness in their attention distributions, which limits robustness in noisy environments. This work proposes a gated graph attention mechanism that employs learnable gates to filter out unreliable features or messages and introduces a learnable temperature parameter to dynamically adjust the sharpness of the attention distribution. By doing so, the method significantly enhances robustness against both feature perturbations and global noise while preserving model expressiveness. Experimental results demonstrate that the proposed model consistently outperforms baseline approaches on both homophilic and heterophilic graph benchmarks and exhibits markedly improved robustness under various noise conditions.
Existing graph Transformers suffer from limited effectiveness, poor scalability, and high preprocessing complexity, often failing to outperform simple GNNs. To address this, we propose the first pure-attention graph learning framework that treats edge sets—not nodes—as the fundamental modeling unit, eliminating conventional node-centric representations and hand-crafted message passing. Our method introduces vertically interleaved masked and standard self-attention encoders, coupled with attention-based pooling for end-to-end differentiable training. It requires no graph reconstruction or preprocessing, natively supports heterogeneous graphs and transfer learning. Evaluated across 70+ node- and graph-level benchmark tasks, our approach consistently surpasses tuned GNN baselines and state-of-the-art graph Transformers. It achieves new SOTA results on molecular graph classification, vision-based graph recognition, heterogeneous graph learning, and cross-domain transfer, while maintaining both high accuracy and linear scalability.
To address the limited modeling capacity of Graph Neural Networks (GNNs) on heterogeneous graphs, this paper proposes Hetero-GAT+LPE, a heterogeneous graph attention network augmented with full-spectrum Laplacian positional encoding. Our method is the first to incorporate full-spectrum Laplacian eigenvectors as positional encodings into heterogeneous graph attention mechanisms, jointly capturing both absolute and relative structural positions of nodes—thereby mitigating the insufficient coupling between semantic and topological information inherent in heterogeneous graphs. Extensive experiments on multiple standard heterogeneous graph benchmarks demonstrate that Hetero-GAT+LPE consistently outperforms state-of-the-art GNNs on node classification and link prediction tasks, achieving average accuracy gains of 3.2%–5.7%. These results empirically validate the critical contribution of structural positional priors to representation learning on heterogeneous graphs.
Graph Transformers (GTs) suffer from diluted local neighborhood information due to global attention, leading to incomplete graph representations. To address this, we propose G2LFormer, the first GT architecture adopting a “global-to-local” attention paradigm: shallow layers capture long-range dependencies, while deeper layers progressively focus on fine-grained local structures; a cross-layer feature fusion mechanism further mitigates representation degradation. G2LFormer integrates linear-complexity graph attention with dedicated GNN modules to enable efficient, synergistic global–local modeling. Evaluated on node and graph classification benchmarks, G2LFormer consistently outperforms state-of-the-art linear-time GTs and classical GNNs. Crucially, it achieves this improvement while maintaining strict O(N) time complexity—enhancing both representational completeness and discriminative power of learned graph embeddings.
Standard Transformers lack the inductive bias necessary for iteratively traversing implicit relational structures in reasoning tasks. To address this limitation, this work proposes the Graph Machine architecture, which introduces an explicit edge mechanism as a novel inductive bias into neural networks. By integrating edge-augmented attention and an edge-centric referencing mechanism, the model enables dynamic, differentiable construction and updating of relational graphs. Evaluated on the Sudoku benchmark, the proposed method significantly outperforms standard Transformers. Ablation studies confirm that the edge mechanism is crucial for performance gains and further reveal that the model automatically learns compact geometric relational representations.
This work addresses the limitations of existing methods in handling heterogeneous node features—such as images and text—in multimodal graphs, where inflexible and inefficient intra- and inter-modal message passing hinders performance. To overcome this, we propose the Dynamic Information Pathway (DiP) framework, which introduces modality-specific pseudo-nodes to construct dynamic, sparse information pathways within a shared state space. This design enables adaptive intra-modal message routing and efficient inter-modal dependency modeling. Notably, DiP achieves cross-modal adaptive propagation with linear complexity, circumventing the constraints of static architectures or dense attention mechanisms. Extensive experiments demonstrate that DiP significantly outperforms state-of-the-art approaches on multiple benchmark datasets for both link prediction and node classification tasks.
Existing graph neural approaches struggle to effectively align modalities and fuse information in multimodal attributed graphs, often neglecting graph context, suppressing cross-modal interactions, and lacking adaptive utilization of topological structure. This work introduces Clifford algebra into multimodal graph learning for the first time and proposes a decoupled propagation-aggregation paradigm: it achieves context-aware high-order modality alignment through modality-aware geometric manifolds and designs an adaptive holographic aggregation mechanism based on geometric-grade energy and scale for effective fusion. The proposed method significantly outperforms state-of-the-art models across nine datasets, achieving leading performance on three types of graph tasks and three types of modality tasks.