heterogeneous graph condensation

Design and build algorithms and pipelines that produce compact, condensed versions of heterogeneous graphs by clustering or summarizing nodes and edges across multiple node/edge types (including methods that explicitly account for differing update or sampling rates). These condensed or abstracted graphs are constructed to preserve the structural and statistical properties required for downstream tasks and model training (e.g., to reduce heterogeneous GNN training time) and to permit reconstruction or mapping back to original graph elements from cluster assignments.

heterogeneousgraphcondensation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Adaptive Graph Coarsening for Efficient GNN Training

Sep 29, 2025
RO
Rostyslav Olshevskyi
🏛️ Rice University

To address the low training efficiency and high memory overhead of Graph Neural Networks (GNNs) on large-scale graphs, this paper proposes an end-to-end learnable adaptive graph coarsening method. Our approach jointly optimizes GNN parameters and node-merging policies during training, employing differentiable K-means clustering on dynamic node embeddings to achieve task-aware graph simplification. Unlike conventional coarsening methods relying on static structural or feature-based heuristics, ours is the first to enable *train-time learnable*, *heterophily-aware* dynamic coarsening. Experiments on both homophilic and heterophilic graph node classification tasks demonstrate that our method significantly reduces computational and memory costs—by up to 5.3×—while preserving or even improving classification accuracy. Visualization further confirms the downstream-task adaptivity of the learned clustering.

Adaptive graph coarsening for efficient GNN trainingEnables graph reduction adaptable to heterophilic and homophilic dataJointly learns GNN parameters and merges nodes during training

This work addresses the challenge of effectively integrating graph structure and node attributes for unsupervised clustering in attributed graphs. The authors propose a multi-round self-training framework based on graph neural networks that alternately refines node representations and cluster assignments in an unsupervised manner. In each round, the current clustering result is used to reconstruct the graph structure, which is then fused with the original graph to form a context-aware graph for generating improved representations. The key innovation lies in the dynamic, synergistic utilization of both edge structure and node attributes, overcoming limitations of conventional single-round training or reliance on a single information source. Experiments demonstrate that the method significantly outperforms baselines using only structure or attributes on synthetic data, that multi-round learning surpasses extended single-round training, and that it achieves state-of-the-art performance on real-world datasets under balanced clustering scenarios.

graph clusteringgraph neural networksnode attributed networks

Rethinking and Accelerating Graph Condensation: A Training-Free Approach with Class Partition

May 22, 2024
XG
Xinyi Gao
🏛️ The University of Queensland | Peking University | Beijing University of Posts and Telecommunications

Existing graph compression (GC) methods rely on bi-level optimization and gradient-based iterations, incurring high computational overhead and slow training. To address this, we propose CGC—a training-free GC framework. Our core insight is to formulate GC as a distribution matching and class partitioning problem between classes and nodes: we introduce the novel “class-to-node” feature matching paradigm, derive a closed-form solution for node features from a predefined graph structure, and directly solve class partitioning via any clustering algorithm—eliminating gradient optimization entirely. On Ogbn-products, CGC compresses the graph in just 30 seconds, achieving 10²–10⁴× speedup over state-of-the-art methods, while boosting downstream GNN accuracy by up to 4.2%. CGC is the first scalable, theoretically tractable, and training-free framework for large-scale graph compression.

Computational EfficiencyFeature MatchingGraph Compression

From Cluster Assumption to Graph Convolution: Graph-Based Semi-Supervised Learning Revisited

Sep 24, 2023
ZW
Zheng Wang
🏛️ Shanghai Jiao Tong University | NIO Technology | University of Macau | University of Illinois at Chicago

This paper addresses the fundamental over-smoothing problem in Graph Convolutional Networks (GCNs), arising from the difficulty of jointly modeling graph structure and label information across layers. We theoretically identify its root cause within a unified optimization framework and clarify the essential divergence between conventional graph-based semi-supervised learning (grounded in the cluster assumption) and GCNs in their optimization objectives. Building on this analysis, we propose three novel graph convolution paradigms: (i) supervised OGC, (ii) learning-free structure-preserving GGC, and (iii) multi-scale GGCM—each explicitly unifying label guidance with structural preservation. Experiments demonstrate that our methods significantly mitigate over-smoothing and consistently outperform mainstream models—including GCN and GAT—on benchmark datasets such as Cora and Citeseer. This validates the critical importance of co-modeling structural fidelity and label supervision.

Addressing GCNs' neglect of graph structure and label informationExploring relationship between traditional and GCN-based semi-supervised learningProposing improved graph convolution methods for better performance

Multi-Scale Node Embeddings for Graph Modeling and Generation

Dec 05, 2024
RM
Riccardo Milocco
🏛️ IMT School for Advanced Studies | ING Bank N.V. | Leiden University

Existing node embedding methods suffer from two key limitations: vector addition lacks network semantic interpretability, and relationships among multi-scale (coarse-grained) embeddings remain ill-defined. This paper proposes a multi-scale node embedding framework that unifies the resolution of both issues for the first time. Leveraging a hierarchical coarse-graining mechanism grounded in renormalization theory, and imposing vector-sum constraints alongside low-dimensional reconstruction optimization in the embedding space, our method ensures that the embedding of any coarse-grained block node is strictly equal to the statistical mean of its constituent node embeddings. This guarantees statistical consistency across resolutions. Evaluated on international trade and input-output networks, the framework achieves high-fidelity structural reconstruction—e.g., accurate triangle counting—and supports arbitrary-scale graph generation. It significantly enhances interpretability and practicality in multi-scale graph modeling and synthesis.

Clarifying the network meaning of vector addition in embeddingsDeveloping consistent multiscale embeddings for network modeling and generationUnderstanding relationships between embeddings at different hierarchical scales

Latest Papers

What's happening recently
View more

Training heterogeneous graph neural networks (HGNNs) on large-scale heterogeneous graphs is computationally expensive, yet existing graph condensation methods struggle to accommodate heterogeneous structures and often rely on costly gradient matching or bi-level optimization. To address this challenge, this work proposes HGC-RC, a novel framework that introduces, for the first time, a role-aware hybrid clustering strategy: it applies class-wise clustering to target nodes while performing type-level unsupervised clustering on non-target nodes. Coupled with lightweight propagation to obtain semantically enriched embeddings, this approach efficiently reconstructs a compact heterogeneous graph. Notably, HGC-RC avoids complex optimization procedures and achieves substantial graph compression while preserving—or even enhancing—downstream task performance, thereby significantly accelerating HGNN training.

computational efficiencygraph condensationheterogeneous graph

This work addresses the challenge of efficiently and accurately clustering temporal graphs to obtain coarse-grained representations when node attributes are missing or weak. The authors propose a theory-driven temporal graph pooling method that formulates community detection as a pooling operator grounded in spectral graph theory. By integrating graph neural networks with multi-slice modularity optimization and incorporating GPU-accelerated spectral clustering and stochastic block models, the approach yields a scalable clustering primitive that remains effective even in attribute-scarce settings. Empirical results demonstrate that while the method excels in scenarios lacking strong attribute signals, neural models achieve superior performance when structural, temporal, and attribute information align coherently. This study thus establishes a new pathway for temporal graph coarsening that balances theoretical rigor with computational efficiency.

clusteringcommunity detectiongraph learning

Graph neural networks (GNNs) often suffer from semantic information loss during pooling operations in graph classification, which hinders their ability to provide interpretability at both subgraph and graph levels. To address this limitation, this work proposes the Subgraph Concept Network (SCN), which employs soft clustering of node concept embeddings to jointly and end-to-end distill semantic concepts at both subgraph and graph granularities. SCN is the first method to enable collaborative learning of multi-level concepts within GNNs, thereby overcoming the conventional reliance on node embeddings alone for interpretation. The approach achieves competitive graph classification performance while significantly enhancing model interpretability through explicit, hierarchical concept discovery.

Concept-based ExplanationsGraph ClassificationGraph Neural Networks

This work addresses the high memory overhead and limited expressiveness of existing methods in large-scale graph neural network (GNN) training, which often fail to simultaneously capture long-range dependencies and preserve node distinguishability. To overcome these limitations, we propose CoRe-GNN, a novel framework that unifies cross-cluster and intra-cluster message passing in parallel: the former leverages graph coarsening to model long-range structural information, while the latter retains fine-grained local node features. By integrating cluster-based batching, CoRe-GNN enables memory-efficient and scalable training with theoretical guarantees on spectral approximation. Empirical results demonstrate that our method consistently outperforms strong baselines such as Cluster-GCN across homophilic, heterophilic, and long-range graphs, achieving state-of-the-art accuracy—particularly in long-range scenarios—while maintaining superior memory efficiency.

Graph CoarseningGraph Neural NetworksLarge-scale Graphs

Hot Scholars

XL

Xunkai Li

School of Computer Science and Technology, Beijing Institution of Technology
Data-centric AIGraph MLAI4Science
CY

Carl Yang

Waymo LLC, PhD at University of California, Davis
GPU ComputingParallel ComputingGraph Processing
YZ

Yinlin Zhu

Sun Yat-sen University
Graph Neural NetworksFederated Learning
MH

Mohammad Hashemi

Emory University
Machine LearningSpatial ComputingGraph Data MiningComputer Vision