clustering-based preprocessing

Design and implement preprocessing pipelines that compute clusters or graph partitions from network-structured data and use those clusters to initialize models, guide partition assignments, or transform the data for streaming processing. These methods focus on preserving global graph structure while maintaining computational efficiency and scalability so as to improve partition quality and downstream graph-model training.

clustering-basedpreprocessing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Streaming Graph Algorithms in the Massively Parallel Computation Model

Jun 17, 2024
AC
Artur Czumaj
🏛️ University of Warwick | National University of Singapore | University of Liverpool

This work addresses dynamic graph processing in the Massively Parallel Computation (MPC) model. We propose the first algorithmic framework supporting batch edge insertions and deletions, terminating in a constant number of parallel rounds, and requiring strongly sublinear local memory per machine. Under the constraint that total memory is sublinear—i.e., significantly smaller than the graph size—we achieve, for the first time, efficient dynamic maintenance of fundamental graph problems, including connectivity, minimum spanning forest, and approximate maximum matching. In contrast to prior approaches relying on linear total memory, our framework attains asymptotic optimality simultaneously in round complexity, local memory usage, and total memory consumption. This breakthrough substantially improves processing efficiency and scalability for massive graphs under resource-constrained environments.

Graph AlgorithmsMemory EfficiencyMPC Model

CluStRE: Streaming Graph Clustering with Multi-Stage Refinement

Feb 08, 2025
AC
Adil Chhabra
🏛️ Heidelberg University

To address the challenge of efficient, high-quality community detection in large-scale dynamic graph streams, this paper proposes a multi-stage refinement streaming graph clustering method. The core innovation lies in (1) constructing and dynamically evolving a quotient graph within the streaming setting for the first time, (2) integrating modularity optimization with a re-streaming mechanism, and (3) introducing an evolution-inspired heuristic alongside a multi-configuration adaptive scheduling strategy. Our method achieves clustering quality exceeding 96% of the offline Louvain algorithm—despite requiring over one-third less memory. Compared to state-of-the-art streaming methods, it improves clustering quality by up to 89.8% (and up to 150% under optimal configuration) while accelerating execution by 2.6×. These advances substantially narrow the performance gap between streaming and batch-mode graph clustering.

Balances efficiency with clustering qualityReduces memory overhead significantlyStreaming graph clustering algorithm

PASCO (PArallel Structured COarsening): an overlay to speed up graph clustering algorithms

Dec 18, 2024
EL
Etienne Lasalle
🏛️ Inria | ENS de Lyon | CNRS | Université Claude Bernard Lyon 1 | LIP | Central European University | National Laboratory for Health Security | HUN-REN Alfréd Rényi Institute of Mathematics

To address the scalability limitations of spectral clustering in large-scale graph clustering, this paper proposes a parallel multi-scale framework. First, a parallelizable structure-preserving graph coarsening algorithm is designed to generate multiple high-fidelity coarse graphs. Second, spectral clustering is executed in parallel on these coarse graphs. Third, optimal transport is introduced—novelly—to align and fuse the resulting multiple partitions, thereby enhancing consistency and clustering quality. By integrating graph coarsening, parallel computation, and optimal transport theory, the method achieves significant speedup while preserving structural fidelity: it delivers several-fold runtime reduction on both synthetic and real-world datasets, while outperforming existing baselines in NMI and F1 scores. Key contributions are: (1) the first structure-preserving coarsening scheme supporting parallel coarsening; and (2) the first application of optimal transport for multi-partition fusion to improve clustering robustness and accuracy.

Accelerates clustering for large graphs with many communitiesCombines partitions using optimal transport for final outputPreserves structural properties during parallel coarsening

MLPrE -- A tool for preprocessing and exploratory data analysis prior to machine learning model construction

Oct 29, 2025
DS
David S Maxwell
🏛️ The University of Texas MD Anderson Cancer Center

To address poor scalability, integration complexity, and inflexible configuration in preprocessing multi-source heterogeneous data for machine learning modeling and graph database construction, this paper proposes a lightweight, modular, JSON-driven automated data preprocessing framework. Built upon Spark DataFrames for efficient distributed processing, the framework defines 69 composable and parallelizable processing stages spanning input parsing, filtering, statistical analysis, feature engineering, and graph-structure transformation. Its declarative JSON-based configuration enables dynamic adaptation to varying data types and scales, significantly enhancing interoperability with workflow orchestration systems such as Apache Airflow. Experimental evaluation across six heterogeneous datasets demonstrates the framework’s generality and scalability: it successfully supports wine quality clustering analysis and end-to-end conversion of phosphosylation site–kinase interaction data into a graph database.

Addressing scalability limitations in existing data processing workflowsPreprocessing diverse data formats for machine learning modelsProviding exploratory analysis capabilities for large-scale datasets

GastCoCo: Graph Storage and Coroutine-Based Prefetch Co-Design for Dynamic Graph Processing

Dec 22, 2023
HL
Hongfu Li
🏛️ Northeastern University | Tongyi Lab | Alibaba Group

Dynamic graph processing faces a fundamental trade-off between computational efficiency—requiring contiguous memory layouts—and update efficiency—necessitating pointer-based, mutable structures—while cache misses dominate performance bottlenecks. To address this, we propose CBList, a prefetch-friendly dynamic graph storage structure, and pioneer the integration of stackless coroutines into graph traversal to enable fine-grained, low-overhead software prefetching. By tightly co-designing CBList’s memory layout with coroutine-driven prefetching, our approach simultaneously achieves high spatial locality and structural mutability. Experimental evaluation demonstrates 1.3×–180× speedup in graph updates and 1.4×–41.1× speedup in graph computation over state-of-the-art dynamic graph systems. Our core contribution is the first systematic co-design paradigm unifying storage structure and coroutine-based prefetching, establishing a novel, cache-aware methodology for dynamic graph processing.

Co-designing storage and prefetch for better graph performanceReducing cache miss overhead in dynamic graph processingResolving conflict between computation and update efficiency in dynamic graphs

Latest Papers

What's happening recently
View more

This work addresses the limitation of existing exact solvers for large-scale Maximum k-Cut problems (k > 2), which stems from the absence of effective preprocessing techniques. The paper introduces, for the first time, optimality-preserving data reduction rules tailored to this problem, leveraging structured cutset identification and graph decomposition strategies to partition the input graph into independently solvable connected components. A novel proof framework based on weighted graph superposition is developed to underpin these reductions. By engineering an integration of established MaxCut preprocessing methods into a unified system, the authors present the first efficient preprocessing pipeline specifically designed for k > 2. Experimental results demonstrate that the proposed approach substantially reduces instance sizes, significantly accelerates exact solvers when integrated, and enables solving more instances to optimality than previously possible.

data reductionexact solversMaximum k-Cut

Existing streaming graph partitioning methods support only vertex or edge partitioning and typically optimize a single objective, making it challenging to simultaneously address the diverse requirements of communication, computation, and memory in distributed GNN training. This work proposes SIGMA, a unified framework that, for the first time in a streaming setting, supports both edge-cut-oriented vertex partitioning and vertex-cut-oriented edge partitioning while jointly optimizing load balancing for both vertices and edges. By incorporating global structural information through a clustering-based preprocessing step, SIGMA significantly enhances partition quality without sacrificing streaming efficiency. Experiments on six large-scale graphs demonstrate that SIGMA outperforms existing streaming methods and achieves partition quality, training efficiency, and memory usage comparable to high-quality offline partitioners such as METIS and KaHIP, while remaining compatible with systems like DistGNN and DistDGL.

distributed GNN trainingedge balancegraph partitioning

Existing hypergraph partitioning methods often become trapped in local optima, limiting partition quality. This work proposes ComPart, a novel framework that integrates community structure guidance during both the initial partitioning and uncoarsening phases. It is the first to comprehensively incorporate community detection throughout the entire uncoarsening process and extends the theory of local dense decomposition from graphs to hypergraphs to generate high-quality initial partitions. By synergistically combining diverse community detection techniques, hypergraph local dense decomposition, and a multilevel partitioning strategy, ComPart consistently outperforms state-of-the-art methods on standard benchmarks, achieving significantly improved partition quality.

community detectionhypergraph partitioninginitial partitioning

This work addresses the challenge of efficiently generating structure-preserving node embeddings for large-scale graphs with millions to billions of edges, which are constrained by memory and computational limitations on single machines. The authors propose an MPI-based distributed graph embedding framework that implements and extends the LINE proximity model. By introducing a communication-computation co-optimization strategy tailored to irregular graph partitions, the framework substantially reduces communication overhead and enhances scalability. Experiments on the NERSC Perlmutter cluster demonstrate that the proposed method achieves 10–100× speedup over multithreaded LINE and node2vec, and 35–76× speedup compared to distributed PyTorch-BigGraph (PBG), while maintaining embedding quality on par with state-of-the-art approaches. End-to-end training time is accelerated by up to 12–370× across various datasets.

distributed graphsgraph embeddingmassive graphs

GPU-Accelerated Algorithms for Process Mapping

Oct 14, 2025
PS
Petr Samoldekin
🏛️ Heidelberg University

This work addresses the classical task graph mapping problem onto processing units in supercomputers, aiming to balance computational load and minimize inter-task communication overhead. For the first time, GPU acceleration is introduced into this domain, yielding two parallel algorithms: (1) a hierarchical multi-partitioning framework accelerated on GPUs, and (2) a GPU-accelerated multilevel graph partitioning implementation integrating optimized coarsening and refinement strategies. Experiments demonstrate speedups of up to 598× over state-of-the-art CPU-based solvers, with a geometric mean speedup of 77.6×; Algorithm (1) incurs only ~10% increase in communication cost while maintaining competitive solution quality. The core contribution is the establishment of a novel GPU-parallel paradigm for task mapping—breaking through long-standing performance bottlenecks inherent in traditional CPU-centric approaches.

GPU-accelerated algorithms balance computational workload and minimize communication costsHierarchical multisection partitions task graphs using supercomputer hierarchyMultilevel graph partitioning pipeline accelerates coarsening and refinement phases

Hot Scholars

HL

Han Lin Shang

Department of Actuarial Studies and Business Analytics, Macquarie University
Functional data analysisnonparametric smoothingnonparametric statisticsmachine learning
SC

SueYeon Chung

Assistant Professor, Harvard University, Flatiron Institite
Theoretical NeuroscienceNeural NetworksStatistical PhysicsMachine Learning
RG

Ruiquan Ge

Hangzhou Dianzi University
Artificial intelligenceBioinformaticsHealth informationImage processing
AE

Ahmed Elazab

PhD, Biomedical engineering
Medical Image AnalysisComputer-aided Detection and DiagnosisMachine & Deep Learningothers
CW

Christoph Weisser

Data Science & Statistics, BASF
Data ScienceEconometricsNatural Language ProcessingDeep Learning