🤖 AI Summary
This study addresses the challenge of graph data scarcity that constrains machine learning development by proposing a graph data augmentation framework based on contrastive generator inversion. Methodologically, the approach infers configuration parameters of the ABCD model through inversion to generate synthetic graphs that preserve macroscopic structural properties. Furthermore, it introduces a multi-positive contrastive objective coupled with a soft negative weighting strategy to enable joint learning of graph representations and generative parameters. Experimental results demonstrate that the proposed method yields substantial improvements in community detection, increasing the Adjusted Mutual Information (AMI) by 161%–273%. Additionally, it achieves significantly superior accuracy and robustness in parameter recovery compared to baseline approaches. This work establishes an effective augmentation paradigm for graph-scarce scenarios.
📝 Abstract
Graphs provide a natural representation of many complex systems, ranging from social platforms to ecosystems. However, the development of graph-based machine learning methods is often constrained by the limited availability of large and diverse graph datasets. In this paper, we introduce $\texttt{DCBA}$, a model-based approach to graph data augmentation that infers the configuration of a synthetic graph generator from an observed network. We instantiate the proposed framework using the $\texttt{ABCD}$ generator, which produces scale-free networks with community structure. Our model learns a joint representation of graphs and generator parametrisations using a multi-positive contrastive objective with soft negative weighting. The learned representation enables the prediction of an $\texttt{ABCD}$ configuration whose stochastic realisations preserve the macrostructural properties encoded by the generator. Experiments show that $\texttt{DCBA}$ recovers generator parameters more accurately and robustly than an algorithmic inverse-modelling baseline. Its downstream utility is further demonstrated in community detection, where inferred configurations used to fine-tune $\texttt{PRoCD}$ improve AMI on average by $161\%$ on synthetic and $273\%$ on real-world networks.