🤖 AI Summary
This study addresses the challenges of joint topology-attribute modeling and scalability in large-scale attributed graph generation by proposing Schema. This model introduces a novel hierarchical decomposition architecture based on soft community memberships, which recursively partitions community structures and trains them independently across stages to sequentially generate node attributes, intra-community edges, and inter-community connections, thereby circumventing full adjacency matrix computation. By integrating graph neural networks with probabilistic graphical models, Schema achieves a balanced generation of local structures and long-range dependencies while generalizing without requiring independent samples. Experimental results demonstrate that Schema maintains high structural fidelity and downstream task accuracy across four real-world datasets, successfully scaling to graphs comprising tens of millions of nodes.
📝 Abstract
Generating large attributed graphs requires reproducing the topology, generating attributes jointly with the structure, and remaining scalable. Many real-world graphs exist as a single large graph, so a generative model has to generalize from the one graph it is fit on, without independent samples. We present Schema, which recursively decomposes a reference graph into a hierarchy of soft communities, assigning each node a membership distribution. Generation is then split into three stages, each trained independently: (1) synthesizing node attributes conditioned on soft memberships, (2) generating intra-community edges from local structural context, and (3) modeling inter-community connections over bridge nodes whose membership mass is distributed across several communities. No stage forms the full adjacency matrix, and each stage operates on a subgraph bounded by the community size. We also introduce an evaluation protocol that covers structural fidelity, memorization, downstream utility, and scalability. On four real-world attributed graphs, Schema recovers the balance between local and long-range structure more closely than any other model that generates attributes, while reproducing only a small fraction of the reference edges. It retains the downstream accuracy of the reference graph without raising it artificially above that level. Baselines that match its structural fidelity memorize the reference, while those with higher downstream accuracy either exceed the reference accuracy or fail to complete on the larger graphs. We measure scalability on six additional graphs with up to 10 million nodes.