Score
Design and build algorithms and pipelines that produce compact, condensed versions of heterogeneous graphs by clustering or summarizing nodes and edges across multiple node/edge types (including methods that explicitly account for differing update or sampling rates). These condensed or abstracted graphs are constructed to preserve the structural and statistical properties required for downstream tasks and model training (e.g., to reduce heterogeneous GNN training time) and to permit reconstruction or mapping back to original graph elements from cluster assignments.
The explosive growth of graph data has imposed severe storage, transmission, and computational bottlenecks on Graph Neural Network (GNN) training, necessitating compact yet high-fidelity condensed graphs for efficient learning. This paper presents a systematic survey of Graph Condensation (GC) techniques. We propose, for the first time, a five-dimensional evaluation framework—encompassing effectiveness, generalizability, efficiency, fairness, and robustness—and unify GC methodologies into two core components: optimization strategies (e.g., gradient matching, feature distillation, meta-learning, adversarial generation) and condensed graph generation (e.g., spectral analysis, topological modeling, differentiable graph synthesis). Through empirical benchmarking of state-of-the-art methods, we analyze open-source ecosystems and cross-domain applications, identifying critical challenges—including limited scalability and lack of support for dynamic graphs—and outline promising future research directions.
To address the low training efficiency and high memory overhead of Graph Neural Networks (GNNs) on large-scale graphs, this paper proposes an end-to-end learnable adaptive graph coarsening method. Our approach jointly optimizes GNN parameters and node-merging policies during training, employing differentiable K-means clustering on dynamic node embeddings to achieve task-aware graph simplification. Unlike conventional coarsening methods relying on static structural or feature-based heuristics, ours is the first to enable *train-time learnable*, *heterophily-aware* dynamic coarsening. Experiments on both homophilic and heterophilic graph node classification tasks demonstrate that our method significantly reduces computational and memory costs—by up to 5.3×—while preserving or even improving classification accuracy. Visualization further confirms the downstream-task adaptivity of the learned clustering.
This work addresses the challenge of effectively integrating graph structure and node attributes for unsupervised clustering in attributed graphs. The authors propose a multi-round self-training framework based on graph neural networks that alternately refines node representations and cluster assignments in an unsupervised manner. In each round, the current clustering result is used to reconstruct the graph structure, which is then fused with the original graph to form a context-aware graph for generating improved representations. The key innovation lies in the dynamic, synergistic utilization of both edge structure and node attributes, overcoming limitations of conventional single-round training or reliance on a single information source. Experiments demonstrate that the method significantly outperforms baselines using only structure or attributes on synthetic data, that multi-round learning surpasses extended single-round training, and that it achieves state-of-the-art performance on real-world datasets under balanced clustering scenarios.
Existing graph compression (GC) methods rely on bi-level optimization and gradient-based iterations, incurring high computational overhead and slow training. To address this, we propose CGC—a training-free GC framework. Our core insight is to formulate GC as a distribution matching and class partitioning problem between classes and nodes: we introduce the novel “class-to-node” feature matching paradigm, derive a closed-form solution for node features from a predefined graph structure, and directly solve class partitioning via any clustering algorithm—eliminating gradient optimization entirely. On Ogbn-products, CGC compresses the graph in just 30 seconds, achieving 10²–10⁴× speedup over state-of-the-art methods, while boosting downstream GNN accuracy by up to 4.2%. CGC is the first scalable, theoretically tractable, and training-free framework for large-scale graph compression.
This paper addresses the fundamental over-smoothing problem in Graph Convolutional Networks (GCNs), arising from the difficulty of jointly modeling graph structure and label information across layers. We theoretically identify its root cause within a unified optimization framework and clarify the essential divergence between conventional graph-based semi-supervised learning (grounded in the cluster assumption) and GCNs in their optimization objectives. Building on this analysis, we propose three novel graph convolution paradigms: (i) supervised OGC, (ii) learning-free structure-preserving GGC, and (iii) multi-scale GGCM—each explicitly unifying label guidance with structural preservation. Experiments demonstrate that our methods significantly mitigate over-smoothing and consistently outperform mainstream models—including GCN and GAT—on benchmark datasets such as Cora and Citeseer. This validates the critical importance of co-modeling structural fidelity and label supervision.
Existing node embedding methods suffer from two key limitations: vector addition lacks network semantic interpretability, and relationships among multi-scale (coarse-grained) embeddings remain ill-defined. This paper proposes a multi-scale node embedding framework that unifies the resolution of both issues for the first time. Leveraging a hierarchical coarse-graining mechanism grounded in renormalization theory, and imposing vector-sum constraints alongside low-dimensional reconstruction optimization in the embedding space, our method ensures that the embedding of any coarse-grained block node is strictly equal to the statistical mean of its constituent node embeddings. This guarantees statistical consistency across resolutions. Evaluated on international trade and input-output networks, the framework achieves high-fidelity structural reconstruction—e.g., accurate triangle counting—and supports arbitrary-scale graph generation. It significantly enhances interpretability and practicality in multi-scale graph modeling and synthesis.
Training heterogeneous graph neural networks (HGNNs) on large-scale heterogeneous graphs is computationally expensive, yet existing graph condensation methods struggle to accommodate heterogeneous structures and often rely on costly gradient matching or bi-level optimization. To address this challenge, this work proposes HGC-RC, a novel framework that introduces, for the first time, a role-aware hybrid clustering strategy: it applies class-wise clustering to target nodes while performing type-level unsupervised clustering on non-target nodes. Coupled with lightweight propagation to obtain semantically enriched embeddings, this approach efficiently reconstructs a compact heterogeneous graph. Notably, HGC-RC avoids complex optimization procedures and achieves substantial graph compression while preserving—or even enhancing—downstream task performance, thereby significantly accelerating HGNN training.
This work addresses the challenge of efficiently and accurately clustering temporal graphs to obtain coarse-grained representations when node attributes are missing or weak. The authors propose a theory-driven temporal graph pooling method that formulates community detection as a pooling operator grounded in spectral graph theory. By integrating graph neural networks with multi-slice modularity optimization and incorporating GPU-accelerated spectral clustering and stochastic block models, the approach yields a scalable clustering primitive that remains effective even in attribute-scarce settings. Empirical results demonstrate that while the method excels in scenarios lacking strong attribute signals, neural models achieve superior performance when structural, temporal, and attribute information align coherently. This study thus establishes a new pathway for temporal graph coarsening that balances theoretical rigor with computational efficiency.
Graph neural networks (GNNs) often suffer from semantic information loss during pooling operations in graph classification, which hinders their ability to provide interpretability at both subgraph and graph levels. To address this limitation, this work proposes the Subgraph Concept Network (SCN), which employs soft clustering of node concept embeddings to jointly and end-to-end distill semantic concepts at both subgraph and graph granularities. SCN is the first method to enable collaborative learning of multi-level concepts within GNNs, thereby overcoming the conventional reliance on node embeddings alone for interpretation. The approach achieves competitive graph classification performance while significantly enhancing model interpretability through explicit, hierarchical concept discovery.
This work addresses the high memory overhead and limited expressiveness of existing methods in large-scale graph neural network (GNN) training, which often fail to simultaneously capture long-range dependencies and preserve node distinguishability. To overcome these limitations, we propose CoRe-GNN, a novel framework that unifies cross-cluster and intra-cluster message passing in parallel: the former leverages graph coarsening to model long-range structural information, while the latter retains fine-grained local node features. By integrating cluster-based batching, CoRe-GNN enables memory-efficient and scalable training with theoretical guarantees on spectral approximation. Empirical results demonstrate that our method consistently outperforms strong baselines such as Cluster-GCN across homophilic, heterophilic, and long-range graphs, achieving state-of-the-art accuracy—particularly in long-range scenarios—while maintaining superior memory efficiency.