Score
Designs, implements, and evaluates self‑supervised graph representation learning systems that mask parts of a graph (nodes, node features, or substructures) and train encoders/decoders to reconstruct the masked node representations or features, yielding node‑ and graph‑level embeddings. Work includes developing adaptive masking strategies, training masked graph autoencoders without discriminators or negative samples, learning embeddings from partial views, and optionally aligning these graph representations to other modalities (e.g., text) or optimizing them for downstream tasks such as retrieval.
Existing surveys on graph foundation models (GFMs) suffer from outdated coverage, ambiguous taxonomies of self-supervised methods, and an overreliance on architecture-specific perspectives—hindering systematic understanding of general graph knowledge learning. To address these limitations, we propose a knowledge-dimensional, three-tiered classification framework (micro–meso–macro), encompassing nine categories of graph knowledge and over 25 pretraining tasks, unifying multi-level representations of nodes, structures, and semantics. We introduce the first knowledge-guided taxonomy for self-supervised GFMs, shifting away from traditional architecture-centric paradigms to accommodate emerging directions such as graph language models. Furthermore, we establish explicit mappings among knowledge types, pretraining tasks, and generalization strategies. This framework comprehensively covers state-of-the-art advances and significantly enhances model interpretability, downstream generalization capability, and cross-task reusability.
Masked feature reconstruction (MFR) in graph self-supervised learning suffers from weak discriminability and conceptual disconnection from contrastive learning. Method: This paper theoretically establishes, under reasonable assumptions, the objective-function equivalence between MFR and node-level graph contrastive learning (GCL). Building on this insight, we propose Contrastive Masked Feature Reconstruction (CMFR)—a unified framework that introduces a novel contrastive reconstruction paradigm: original and reconstructed features serve as positive pairs, while masked nodes act as negatives. CMFR integrates a context-aware encoder and a customized negative sampling strategy. Contribution/Results: On multiple benchmark datasets, CMFR consistently outperforms GraphMAE and GraphMAE2, achieving up to 3.82% absolute improvement in both node and graph classification tasks, setting a new state-of-the-art in graph self-supervised learning.
This work addresses two key challenges in graph neural networks (GNNs): the difficulty of identifying latent communities and the ambiguity of node classification boundaries. To this end, we propose an encoder-based embedding refinement method that jointly integrates linear transformation, self-training, and implicit community recovery—uniquely coupling these components within the encoder optimization pipeline for the first time. Under the stochastic block model (SBM), we provide theoretical guarantees on convergence and improved community identifiability. Extensive experiments on both synthetic and real-world graph datasets demonstrate that our method significantly enhances latent community detection accuracy and boosts node classification performance by 3.2–7.8% on average. Moreover, it improves model robustness and generalization. The core contributions are: (i) a novel synergistic embedding refinement framework unifying implicit community discovery and self-training; and (ii) theoretically grounded enhancement of community identifiability in GNNs.
Existing research on masked image modeling (MIM) for self-supervised visual representation learning lacks a unified formalization of pretraining paradigms and standardized, comparable evaluation. Method: We formally categorize MIM into two principal paradigms—reconstruction-based and contrastive-based—and construct an interpretable, hierarchical taxonomy via expert curation and agglomerative clustering. We conduct systematic, unified benchmarking of over 20 state-of-the-art models on ImageNet and other standard datasets. Contribution/Results: Our analysis identifies critical open challenges—including cross-paradigm integration, long-tailed masking strategies, and compute-accuracy trade-offs. To foster reproducibility and standardization, we publicly release a structured literature repository and an extensible evaluation framework on GitHub. This work establishes foundational infrastructure for rigorous, comparable advancement in MIM research.
To address the limitations of contrastive learning—such as susceptibility to overfitting and difficulty in capturing semantic hierarchies—in graph-level representation learning, this paper proposes Graph-JEPA, the first adaptation of the Joint Embedding Predictive Architecture (JEPA) to graph-structured data. Graph-JEPA enables contrastive-free and reconstruction-free self-supervision by masking subgraphs and predicting their latent representations. Crucially, it introduces hyperbolic coordinate regression as a novel objective to explicitly model the implicit hierarchical structure among graph concepts. By eliminating negative sampling and pixel-level reconstruction, Graph-JEPA significantly mitigates overfitting. Extensive experiments demonstrate that Graph-JEPA consistently outperforms state-of-the-art self-supervised methods on graph classification, continuous-value regression, and non-isomorphic graph discrimination tasks. The learned graph-level representations exhibit superior semantic richness and generalization capability.
Existing graph Transformers suffer from limited effectiveness, poor scalability, and high preprocessing complexity, often failing to outperform simple GNNs. To address this, we propose the first pure-attention graph learning framework that treats edge sets—not nodes—as the fundamental modeling unit, eliminating conventional node-centric representations and hand-crafted message passing. Our method introduces vertically interleaved masked and standard self-attention encoders, coupled with attention-based pooling for end-to-end differentiable training. It requires no graph reconstruction or preprocessing, natively supports heterogeneous graphs and transfer learning. Evaluated across 70+ node- and graph-level benchmark tasks, our approach consistently surpasses tuned GNN baselines and state-of-the-art graph Transformers. It achieves new SOTA results on molecular graph classification, vision-based graph recognition, heterogeneous graph learning, and cross-domain transfer, while maintaining both high accuracy and linear scalability.
Existing masked graph autoencoders exhibit poor generalization on heterophilic graphs, primarily due to overreliance on the homophily assumption—i.e., the expectation that neighboring nodes share similar features and labels—while neglecting intrinsic node dissimilarities, leading to representation ambiguity. This work pioneers the integration of node dissimilarity modeling into the masked graph autoencoding framework, proposing a dissimilarity-aware neighborhood reconstruction mechanism: during encoding, it explicitly captures and reconstructs feature- and structure-level discrepancies among neighboring nodes, thereby enhancing the discriminability of low-dimensional representations. The method requires no auxiliary labels or external pretraining. Evaluated across 17 benchmark datasets, it consistently outperforms state-of-the-art approaches on node classification, clustering, and graph classification tasks, achieving new best results and substantially advancing heterophilic graph representation learning.
This study investigates how mask design affects downstream performance in molecular graph self-supervised learning. We propose a unified probabilistic framework to quantify the information content of pretraining signals and, under rigorously controlled experimental conditions, disentangle the individual contributions of mask distribution, prediction objective, and encoder architecture. Empirical results reveal that mask distribution (e.g., uniform vs. structure-aware sampling) has limited impact; instead, the semantic richness of the prediction objective—and its synergy with Graph Transformer architectures—determines model performance. Furthermore, node-level masking combined with information-theoretic measures enhances model interpretability. We thus establish the first interpretable, analysis-oriented framework for mask design in molecular graphs, providing both theoretical foundations and practical guidelines for self-supervised molecular representation learning. (132 words)
This work addresses the misalignment between graph representations and the latent feature space of frozen large language models in graph retrieval-augmented generation. To bridge this gap, the authors propose AGE, a method built upon a masked self-supervised learning framework that encodes graphs using a Transformer architecture with text-like embeddings. A key innovation of AGE is its learnable node sampler, which adaptively identifies and bypasses hard-to-predict yet critical nodes to better align graph and textual embedding spaces. Evaluated on four heterogeneous GraphQA benchmark datasets, AGE substantially outperforms existing non-parametric retrieval approaches, achieving state-of-the-art accuracy.
This work addresses the susceptibility of existing graph self-supervised learning methods to low-level input statistics and their limited capacity to model structural relationships among nodes. It introduces, for the first time, the Joint-Embedding Predictive Architecture (JEPA) paradigm to node-level graph representation learning through a structure-conditioned prediction mechanism: by masking k-hop subgraph structures, a context encoder predicts the target node’s representation in latent space, thereby circumventing reliance on input reconstruction or handcrafted augmentations. The approach integrates an EMA target encoder, cross-attention over spectral and centrality descriptors, and variance/covariance/Laplacian regularization, complemented by a progressive curriculum masking strategy to explicitly reinforce structural information learning. Evaluated on standard node classification benchmarks, the method achieves strong performance under both linear probing and fine-tuning, with ablation studies confirming the contribution of each component.