Learning the Latent Structure: A Feature-Centric Approach to Graph Data Augmentation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses structural incompleteness in real-world graph data caused by collection limitations, as well as the reliance of existing augmentation methods on labels, their computational expense, and restriction to transductive settings. To overcome these challenges, we propose SelfAug, a framework that introduces a novel feature-centric augmentation paradigm operating directly in the embedding space to bypass explicit structural modeling. Specifically, it employs a self-supervised inverse masking mechanism to recover unobserved structural signals, complemented by message regularization and bootstrapping strategies to enhance robustness. Extensive experiments across ten multi-domain datasets demonstrate that SelfAug surpasses state-of-the-art methods in both accuracy and efficiency. Furthermore, it achieves effective generalization in inductive learning and cold-start scenarios, validating its potential as a scalable and versatile solution for graph representation learning.
📝 Abstract
Graph-structured data plays a pivotal role in modeling complex relationships. However, real-world graphs are often incomplete due to data collection and observational constraints, severely limiting the effectiveness of modern graph learning pipelines. While existing Graph Data Augmentation (GDA) methods attempt to refine graph structures for improved downstream performance, they are typically label-dependent, computationally expensive, and inherently transductive, limiting their applicability in practical scenarios. In this work, we present a novel feature-centric graph data augmentation framework that bypasses explicit structure modeling by operating directly in the embedding space. Through a self-supervised inverse masking process, our method captures latent ties between observed and complete graphs, enabling recovery of unobserved structural signals through refined node representations. To enhance robustness under noisy and sparse supervision, we introduce a message regularizer and a bootstrap strategy for effective training and generalization. Evaluated on ten graph datasets spanning multiple domains, our approach, SelfAug, consistently outperforms state-of-the-art methods in both accuracy and efficiency across inductive and cold-start settings, highlighting its potential as a scalable and generalizable solution for real-world graph learning scenarios.
Problem

Research questions and friction points this paper is trying to address.

Graph Data Augmentation
Incomplete Graphs
Label Dependency
Inductive Learning
Graph Representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph Data Augmentation
Feature-Centric
Self-Supervised Inverse Masking
Message Regularizer
Inductive Learning
🔎 Similar Papers
No similar papers found.