self-supervised graph learning

Design and implement self-supervised objectives, data augmentations, and neural architectures that learn node-, edge-, subgraph-, or whole-graph embeddings from graph topology and attributes without labeled targets. Build and analyze these representations to enable downstream analyses (e.g., classification, link prediction, clustering, anomaly detection) and to support adaptation or transfer without full retraining.

self-supervisedgraphlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.47
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Leveraging Joint Predictive Embedding and Bayesian Inference in Graph Self Supervised Learning

Feb 02, 2025
SS
Srinitish Srinivasan
🏛️ Vellore Institute of Technology

Existing graph self-supervised learning methods suffer from high computational overhead, reliance on contrastive loss and negative sampling, susceptibility to representation collapse, and difficulty in quantifying the semantic contribution of node embeddings to downstream tasks. To address these issues, we propose a contrastive-free, negative-sampling-free joint embedding prediction framework. Our approach introduces a subgraph-level single-context–multi-target joint prediction mechanism that jointly leverages structural and semantic information; employs a Gaussian Mixture Model (GMM)-driven semantic contribution scoring strategy to generate high-quality pseudo-labels; and incorporates Bayesian inference for robust self-training. Evaluated on multiple benchmark datasets, the framework achieves significant improvements in node classification and link prediction performance, while exhibiting higher training efficiency and effectively mitigating representation collapse.

Eliminates contrastive objectives and negative samplingEnhances node discriminability with pseudo-labelsImproves graph self-supervised learning efficiency

A Hybrid Supervised and Self-Supervised Graph Neural Network for Edge-Centric Applications

Jan 21, 2025
EB
Eugenio Borzone
🏛️ Research institute for signals, systems and computational intelligence (sinc(i))

This work addresses edge-centric tasks—such as protein–protein interaction prediction and similarity assessment of structurally unknown compounds—by proposing a graph neural network framework that unifies supervised and self-supervised learning. Methodologically, it is the first to integrate both loss terms into a unified edge-level prediction objective; introduces an edge-aware attention mechanism that jointly models node and edge features; and incorporates self-supervised contrastive learning with a lightweight feed-forward prediction head, enabling end-to-end edge representation learning using only one-hot node features. Key contributions include: (1) resolving the long-standing challenge of similarity prediction for compounds lacking 3D structural information; and (2) achieving state-of-the-art performance on both protein–protein interaction prediction and gene ontology functional annotation tasks.

Entity Interaction PredictionProtein-Protein InteractionUnknown Compound Similarity

This paper addresses self-supervised representation learning for graph data with no or few labels, proposing the first unified framework that jointly integrates generative and contrastive paradigms. Methodologically, it introduces a community-aware joint node- and graph-level contrastive learning mechanism; employs multi-granularity graph augmentations—including feature masking and node/edge perturbations—to construct robust views; and pioneers end-to-end co-optimization of generative loss (graph reconstruction) and contrastive loss (semantic alignment). The key innovation lies in explicitly embedding community structure priors into the contrastive objective and enabling seamless joint optimization of both learning objectives. Extensive experiments on multiple benchmark datasets demonstrate consistent superiority over state-of-the-art methods: improvements of 0.23%–2.01% are achieved across node classification, clustering, and link prediction tasks.

Enhances node and graph representation robustness and diversityImproves performance in node classification, clustering, and link predictionIntegrates generative and contrastive learning for graph SSL

Improving Graph Neural Networks on Multi-node Tasks with the Labeling Trick

Apr 20, 2023
XW
Xiyuan Wang
🏛️ Peking University | Georgia Institute of Technology

Existing graph neural networks (GNNs) primarily learn single-node representations; direct aggregation fails to capture intra-set dependencies within multi-node structures—such as links, hyperedges, or subgraphs. Method: We propose the “labeling trick”: pre-labeling the target node set, reconstructing the graph structure accordingly, and then encoding and aggregating via standard GNNs—enabling end-to-end learning of multi-node representations. Contribution/Results: We formally characterize this paradigm for the first time, theoretically demonstrating its capacity to model higher-order node dependencies. The framework unifies and generalizes to partially ordered sets, subsets, and hypergraphs. Compatible with mainstream architectures (e.g., GCN, GAT), it supports undirected/directed link prediction, hyperedge prediction, and subgraph classification. On multiple benchmarks, it achieves average improvements of 3.2% in AUC (link prediction), 5.7% in F1-score (hyperedge prediction), and 4.9% in accuracy (subgraph classification), validating both effectiveness and broad applicability.

Addressing multi-node representation learning limitations in GNNsExtending method to graphs, posets, subsets and hypergraphsProposing labeling trick to capture node set dependencies

Latest Papers

What's happening recently
View more

This work addresses the challenge of effectively integrating graph structure and node attributes for unsupervised clustering in attributed graphs. The authors propose a multi-round self-training framework based on graph neural networks that alternately refines node representations and cluster assignments in an unsupervised manner. In each round, the current clustering result is used to reconstruct the graph structure, which is then fused with the original graph to form a context-aware graph for generating improved representations. The key innovation lies in the dynamic, synergistic utilization of both edge structure and node attributes, overcoming limitations of conventional single-round training or reliance on a single information source. Experiments demonstrate that the method significantly outperforms baselines using only structure or attributes on synthetic data, that multi-round learning surpasses extended single-round training, and that it achieves state-of-the-art performance on real-world datasets under balanced clustering scenarios.

graph clusteringgraph neural networksnode attributed networks

Existing self-supervised learning methods exhibit limited performance on link prediction tasks in graphs without node attributes. This work proposes the first self-supervised learning framework centered on link representations, introducing a link-level contrastive learning mechanism that integrates instance discrimination with a community structure-aware graph augmentation strategy. The proposed models, L-GRACE and L-BGRL, significantly outperform current state-of-the-art approaches under both self-supervised and supervised settings, achieving particularly strong results on attribute-free graphs. These empirical gains validate the effectiveness of link-centric representation learning and structure-aware augmentation for improving link prediction performance in the absence of node features.

graph representationinstance discriminationlink prediction

This work addresses the limited generalization of existing node representation learning methods, which typically require dataset-specific training and hyperparameter tuning. The authors propose Node4All, the first framework enabling universal node representation learning across arbitrary graphs with a single model. Built upon a Channel Graph Transformer (CGT) architecture and powered by synthetic graph-driven self-supervised learning, Node4All eliminates the need for dataset-specific fine-tuning and supports both zero-shot and in-context learning. Evaluated on 25 benchmark datasets, it achieves an average rank of 5th, significantly outperforming most baselines and surpassing current graph foundation models under one-shot and in-context settings, thereby overcoming the dataset-specific limitations of conventional approaches.

dataset-specific optimizationgraph generalizationgraph models

Graph self-supervised learning suffers from high computational costs and data redundancy due to its reliance on large-scale unlabeled data. This work proposes the first label-free coreset construction method for graphs, which integrates intrinsic structural diversity with contextual semantics generated by language models. By leveraging graph statistical features, graph-to-text generation, pretrained language model embeddings, and cluster-aware sampling, the approach achieves highly efficient data compression. Remarkably, using only 10% of the original data, the method attains 99.6% of the full-data performance while reducing pretraining time by nearly 90%. Theoretical analysis further provides a guaranteed bound on the approximation loss, ensuring robustness and reliability of the compressed representation.

computational efficiencydata redundancygraph representation

Existing graph self-supervised learning methods are typically confined to a single level of abstraction and employ uniform penalty strengths, limiting their ability to flexibly integrate multi-granularity structural information. This work proposes a unified multi-level contrastive learning framework that simultaneously models representations at the node, neighborhood, cluster, and graph levels, optimizing them through a linear combination of similarity and dissimilarity scores. Additionally, a parameter-free, fine-grained self-weighting mechanism is introduced to dynamically adjust sample weights during training, enhancing optimization efficiency. To the best of our knowledge, this is the first approach to achieve unified modeling of multi-level graph representations, consistently outperforming state-of-the-art methods across node classification, clustering, and link prediction tasks on multiple real-world datasets.

Contrastive LearningGraph RepresentationGraph Self-Supervised Learning

Hot Scholars

CM

Chee-Ming Ting

Associate Professor, Monash University. PhD (Maths - Statistics)
Statistical Signal ProcessingMachine LearningBiomedical SignalsNeuroimaging
ZY

Ze Yang Ding

Monash University Malaysia
Industrial AIDeep LearningData-Driven ModelingSoft Sensor
MF

Marcelo Fiori

PhD, Universidad de la República, Uruguay
Graph Matching ProblemsSignal Processing
WZ

Wenjie Zhang

Professor of Computer Science and Engineering, University of New South Wales
database systemsbig data analyticsdata-centric AI
FG

Fausto Giunchiglia

Professor of Computer Science, Università di Trento
Computational theories of the mind