Score
Designs and implements algorithms and models that generate graph-structured data samples that are distinct from existing graphs while preserving desired global structural properties, often via latent-space embeddings or other generative mechanisms. Develops metrics, assessment procedures, and risk-quantification methods to measure, control, and produce novelty-aware graph generation and evaluate the novelty of produced graphs.
This work addresses the challenge of generating significantly novel graph samples while preserving global structural consistency. To this end, the authors propose a novel approach based on latent space embeddings and a constrained mixture model, which—by explicitly modeling novelty and fidelity constraints—enables theoretically grounded, controllable graph generation. Notably, this is the first study to incorporate the Minimum Description Length (MDL) principle from information theory into graph generation, providing formal guarantees that the probability of erroneously accepting non-novel or unreliable samples decays to zero at an explicit rate as the threshold tightens. Extensive experiments on both synthetic and standard benchmark graph datasets demonstrate the method’s effectiveness, achieving principled and quantifiably low-risk generation of novel graphs.
Real-world network flow data is often inaccessible due to privacy, security, and computational constraints. Method: This paper proposes a high-fidelity, diversity-controllable dynamic multigraph synthesis framework. Its core innovation lies in the first joint optimization of structural accuracy and attribute diversity. It employs a decoupled modeling strategy: Kronecker graph generation for topology, Tabular GANs for node/edge attribute synthesis, and an XGBoost-driven graph alignment mechanism to coordinate structural and attribute optimization. Customized evaluation metrics are designed to quantify synthesis quality. Results: Experiments on large-scale NetFlow data demonstrate that the method significantly outperforms existing graph generation techniques, achieving superior trade-offs among fidelity, diversity, and computational efficiency. Multiple validated synthetic datasets—capable of supporting privacy-sensitive modeling tasks—are successfully generated and empirically verified.
Existing random graph models (e.g., Erdős–Rényi, Kronecker) assume edge independence, limiting their ability to simultaneously achieve high subgraph density (e.g., triangles), high output variability, and realistic structural properties—such as power-law degree distributions, high clustering, and small diameter. To address this, we propose a novel edge-dependent graph generation framework grounded in a “binding” mechanism. This work establishes, for the first time, a provably sound theoretical foundation for preserving output variability under edge dependence and derives closed-form expressions for subgraph densities. The method offers high controllability, computational efficiency (linear-time complexity), and enhanced graph diversity. Experimental results demonstrate that graphs generated by our approach exhibit a 2.3× higher clustering coefficient and a 47% increase in coefficient of variation compared to baseline models, significantly improving structural fidelity to real-world networks.
Graph anomaly detection (GAD) is critical for security, finance, and other domains, yet existing GNN-based approaches lack systematic organization and a unified analytical framework. To address this, we propose the first comprehensive analysis paradigm grounded in three orthogonal dimensions: GNN backbone design, proxy task construction, and anomaly scoring. We introduce a fine-grained taxonomy comprising 13 categories, decoupling model architecture into backbone networks, pretraining objectives, and anomaly criteria. Integrating GNNs, self-supervised learning, contrastive learning, reconstruction modeling, and multi-scale representation, we establish a reproducible benchmarking suite. Our open-source, continuously updated repository unifies state-of-the-art datasets and algorithms, accompanied by empirical performance comparisons. The study exposes intrinsic limitations of current methods and identifies six key open challenges—providing both theoretical guidance and practical foundations for future GAD research.
Conventional diffusion models for graph generation suffer from O(n²) computational complexity in the node space, hindering scalability. Method: This paper proposes GGSD, the first model to jointly integrate graph Laplacian spectral decomposition with denoising diffusion probabilistic modeling, establishing a novel spectral-space diffusion paradigm. GGSD employs spectral truncation for efficient low-dimensional representation and introduces a linear-complexity, permutation-invariant Transformer architecture that supports node feature fusion. Crucially, it generates graph structures directly in the node space while achieving theoretical O(n) complexity. Contribution/Results: Extensive experiments demonstrate that GGSD significantly outperforms state-of-the-art methods on both synthetic and real-world graph datasets, achieving superior trade-offs among generation speed, structural fidelity, and scalability.
This work addresses the challenge of generating realistic graph data while preserving critical structural properties such as degree distribution and spectral characteristics, which are often distorted in existing methods. The authors propose a novel hybrid approach that integrates Wasserstein GAN (WGAN) with a genetic algorithm (GA). Initially, WGAN produces a coarse graph, which is subsequently refined through GA-based evolutionary optimization of edge connections in a gradient-free manner. Notably, this is the first study to incorporate GA into the post-processing stage of GAN-based graph generation, using Maximum Mean Discrepancy (MMD) as the optimization objective. The method achieves a significant reduction in aggregate MMD metrics, yielding synthetic graphs that closely match real-world data in key topological features while maintaining diversity, thereby enhancing both the quality and practical utility of generated graphs.
Existing graph generation models struggle to balance scalability and novelty, often failing to efficiently produce realistic and diverse graph structures. This work proposes a lightweight autoregressive framework that serializes graphs into edge sequences via structure-guided topological ordering and employs a two-stage exploration–refinement training strategy to reduce computational complexity while enhancing generalization and controllable novelty. The approach is compatible with sequential architectures such as LSTM and Mamba and incorporates a large-memory acceleration technique to overcome GPU memory constraints. Experimental results demonstrate that, on both molecular and non-molecular benchmarks, the generated graphs achieve high validity and uniqueness while significantly improving novelty and diversity.
Existing network generation methods often suffer from overfitting, neglect of critical structural features, and high computational costs, hindering the efficient synthesis of high-fidelity networks. To address these limitations, this work proposes SyNGLER, a framework that learns node embeddings in a low-dimensional latent space and employs a distribution-agnostic generator to resample and reconstruct networks within this space. The approach effectively preserves essential topological properties—such as sparsity and degree heterogeneity—and provides theoretical guarantees on edge distribution consistency. Experimental results demonstrate that SyNGLER significantly reduces computational overhead while more accurately reproducing the degree distributions and higher-order moments of real-world networks, outperforming current deep generative models.