Cluster Analysis with Resampling for Validation and Exploration (CARVE)

๐Ÿ“… 2026-05-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Clustering outcomes are highly sensitive to algorithmic choices, preprocessing steps, and the number of clusters, yet conventional validation metrics often fail in high-dimensional, heavy-tailed, or nonlinear biomedical data, leading to irreproducible findings. This work proposes a resampling-driven framework for clustering evaluation that unifies stability and generalization analyses for the first time, enabling diagnostic assessments at global, cluster-level, and sample-level resolutions while producing consensus cluster labels and selection criteria. By circumventing restrictive geometric assumptions inherent in traditional methods, the approach offers a scikit-learnโ€“compatible Python API and a Seurat-compatible R interface. It consistently approximates optimal clustering across six synthetic benchmarks, significantly outperforming existing metrics, and uncovers finer biological structures in real-world genomics and proteomics datasets.
๐Ÿ“ Abstract
Clustering is widely used across the sciences as the foundation for downstream data-driven scientific discoveries. However, clustering results are highly sensitive to the choice of algorithm, preprocessing, and the number of clusters $k$, producing scientific claims that are often not reproducible. The current state of the art for validating clustering solutions consists of clustering validation indices (CVIs) such as Silhouette, Davies-Bouldin, and Calinski-Harabasz, which rely on geometric assumptions that break down on the heavy-tailed, high-dimensional, and nonlinearly structured data encountered in biomedical research. Resampling-based alternatives - grounded in the ideas of clustering stability and generalizability - have been proposed but remain scattered across specialized tools with no unified, accessible software. We fill this gap with CARVE (Cluster Analysis with Resampling for Validation and Exploration), an open-source Python and R package that jointly evaluates multiple clustering algorithms and hyperparameters, returning stability and generalizability diagnostics at the global, cluster, and sample level together with principled selection rules and consensus-based cluster labels. Across six synthetic benchmarks CARVE consistently recovers near-optimal clusterings where classical indices degrade substantially. On experimental genomics and proteomics data sets, CARVE recovers finer biological structure when classical CVIs collapse entirely. CARVE is available with a scikit-learn-compatible Python API and an analogous R interface compatible with Seurat workflows.
Problem

Research questions and friction points this paper is trying to address.

clustering validation
resampling
reproducibility
high-dimensional data
cluster stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

resampling
clustering validation
stability
generalizability
consensus clustering
๐Ÿ”Ž Similar Papers
2024-09-01arXiv.orgCitations: 4
K
Kai R. Wycik
Department of Statistics, Columbia University, New York, NY, USA
Tiffany M. Tang
Tiffany M. Tang
Department of Applied and Computational Mathematics and Statistics, University of Notre Dame
statistical machine learningdata sciencegenomics
T
Tarek M. Zikry
School of Data and Information Sciences, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA
Genevera I. Allen
Genevera I. Allen
Columbia University
statistical machine learningneuroscience