Score
Design and implement graph neural network architectures and learning algorithms that incorporate explicit region-level structure—attention mechanisms and pooling operations that are aware of node-to-region assignments—to compute region-preserving node, edge, and graph representations. Use these components to analyze or predict structured system behavior and to enable interpretable, region-consistent outputs and transfer of learned models across different region layouts.
Existing deep learning models struggle to effectively encode spatial, topological, and semantic structural information inherent in images. This work systematically evaluates the impact of various visual graph construction strategies on image classification performance within a unified three-layer Graph Convolutional Network (GCN) framework. For the first time, it demonstrates that the graph structure itself plays a decisive role in model performance. The study underscores the critical importance of the graph construction preprocessing stage, providing empirical evidence that well-designed graph structures substantially enhance classification accuracy. These findings offer both methodological guidance and practical justification for graph structure selection and preprocessing in visual graph neural networks.
Graph neural networks (GNNs) remain challenging to understand and apply for secondary-school students and machine learning practitioners due to conceptual abstraction and fragmented pedagogical resources. Method: We propose a unified encoder–decoder pedagogical and practical framework that systematically integrates core GNN mechanisms—including message passing, GCN, and GAT—and designs task-specific decoders for node classification, link prediction, and other downstream tasks. Grounded in engineering practice, we develop the first reproducible GNN入门 (introductory) curriculum, combining theoretical exposition with large-scale homogeneous-graph experiments. Contribution/Results: We quantitatively characterize how model performance scales with training data size and graph structural complexity—revealing novel empirical patterns. The framework provides standardized benchmarking protocols, hyperparameter tuning guidelines, and task-adaptation strategies. It significantly enhances pedagogical interpretability and industrial deployability of GNNs, lowering barriers to entry without sacrificing technical rigor.
This work identifies inherent limitations of Graph Neural Networks (GNNs) in global structural reasoning—including prototype identification, symmetry detection, connectivity assessment, and critical node localization—as well as in scale-invariant inference. To address these gaps, we introduce GraphAbstract, the first benchmark explicitly designed for global topological awareness, and propose a novel paradigm for graph structural understanding grounded in pretrained vision models: graphs are encoded into structure-preserving visual representations, enabling zero-shot or few-shot visual transfer for cross-scale generalization. Empirical results demonstrate that vision models significantly outperform state-of-the-art GNNs across multiple global property recognition tasks, validating their capacity to capture long-range dependencies and holistic topology without explicit message passing. This approach offers a scalable, robust, and interpretable alternative to conventional graph learning.
Existing graph Transformers suffer from limited effectiveness, poor scalability, and high preprocessing complexity, often failing to outperform simple GNNs. To address this, we propose the first pure-attention graph learning framework that treats edge sets—not nodes—as the fundamental modeling unit, eliminating conventional node-centric representations and hand-crafted message passing. Our method introduces vertically interleaved masked and standard self-attention encoders, coupled with attention-based pooling for end-to-end differentiable training. It requires no graph reconstruction or preprocessing, natively supports heterogeneous graphs and transfer learning. Evaluated across 70+ node- and graph-level benchmark tasks, our approach consistently surpasses tuned GNN baselines and state-of-the-art graph Transformers. It achieves new SOTA results on molecular graph classification, vision-based graph recognition, heterogeneous graph learning, and cross-domain transfer, while maintaining both high accuracy and linear scalability.
Existing studies lack quantitative theoretical analysis of the differences in optimization and generalization performance between graph neural networks (GNNs) and multilayer perceptrons (MLPs). Method: Leveraging feature learning theory, this paper provides the first rigorous analysis of two-layer graph convolutional networks (GCNs) trained via gradient descent, modeling ReLU^q activation functions and spectral properties of the expected degree matrix D, under a structure-guided signal-noise separation framework. Contribution/Results: We theoretically prove that graph convolution explicitly exploits graph structure to significantly enhance signal learning while suppressing noise memorization, thereby expanding the benign overfitting regime by approximately √D^{q−2} compared to CNNs. This constitutes the first quantitative characterization of the fundamental generalization advantage of GNNs over MLPs. Extensive simulations corroborate the superior generalization and robustness of GCNs predicted by our theory.
This work addresses a key limitation of conventional graph neural networks (GNNs), which typically rely solely on the final-layer node representations for pooling or classification, thereby discarding valuable historical activation information from intermediate layers and suffering from issues such as oversmoothing and representational degradation. To overcome this, the authors propose HISTOGRAPH, a two-stage attention-based aggregation framework that first unifies intermediate activations across all GNN layers through inter-layer attention and then dynamically models the evolution of node representations across depths via node-level attention. By systematically leveraging historical activation signals, HISTOGRAPH effectively mitigates information loss in deep GNNs, achieving state-of-the-art performance on multiple graph classification benchmarks and demonstrating enhanced expressiveness and robustness, particularly in deep architectures.
This work addresses a critical yet overlooked issue in graph neural networks (GNNs): during link prediction training, mini-batch sampling can introduce class composition bias, which—particularly when combined with batch normalization—leads models to learn spurious heuristics. Consequently, the learned representations become misaligned with those beneficial for node classification, challenging the common assumption that GNNs produce task-agnostic, transferable embeddings. The study systematically uncovers this bias for the first time and proposes a correction mechanism that aligns model representations with the intrinsic structural properties of the graph. Experimental results demonstrate that the corrected models focus more on class-relevant features, suggesting that standard training protocols may substantially overestimate the generalization capability of link predictors.
Graph neural networks (GNNs) often suffer from semantic information loss during pooling operations in graph classification, which hinders their ability to provide interpretability at both subgraph and graph levels. To address this limitation, this work proposes the Subgraph Concept Network (SCN), which employs soft clustering of node concept embeddings to jointly and end-to-end distill semantic concepts at both subgraph and graph granularities. SCN is the first method to enable collaborative learning of multi-level concepts within GNNs, thereby overcoming the conventional reliance on node embeddings alone for interpretation. The approach achieves competitive graph classification performance while significantly enhancing model interpretability through explicit, hierarchical concept discovery.
Graph pooling often struggles to consistently surpass the expressive power of 1-WL-equivalent GNNs in graph classification tasks, yielding limited performance gains. This work identifies that effective pooling hinges on the alignment between node features and graph topology, and for the first time formally defines the fundamental conditions that node features must satisfy to enable such effective pooling. Building on this insight, we introduce a quantitative metric to measure the degree of feature–topology alignment. Both theoretical analysis and empirical experiments demonstrate that when node features meet the proposed conditions, graph pooling can significantly enhance classification performance on suitable datasets.
This study addresses the lack of a systematic survey on the integration of attention mechanisms with graph neural networks (GNNs), which has hindered a clear understanding of their developmental trajectory. To bridge this gap, the work proposes a novel two-level taxonomic framework that organizes the field both historically and architecturally. At the upper level, it delineates three chronological phases: Graph Recurrent Attention Networks, Graph Attention Networks, and Graph Transformers. The lower level systematically catalogs representative models within each phase and compares their key characteristics. Through comprehensive literature review and taxonomy-based analysis, the paper elucidates the evolutionary pathway of attention in GNNs, clarifies the strengths and limitations of existing approaches, identifies open challenges, and outlines promising future directions. An accompanying open-source repository is provided to foster ongoing community research.