π€ AI Summary
This work addresses the challenge of weak signal strength or limited sample size in high-dimensional clustering by studying transfer clustering under a two-community Gaussian mixture model. It characterizes the relatedness between target and source data through geometric alignment of their cluster means and proposes an adaptive transfer clustering method that achieves minimax optimality. The study establishes, for the first time in high dimensions, a sharp phase transition threshold for transfer clustering consistency, revealing a precise boundary determined jointly by signal strength, sample size, ambient dimension, and alignment quality. The framework naturally extends to multi-community and multi-source settings. Numerical experiments and real-data analysis on human lung single-cell RNA sequencing demonstrate the methodβs effectiveness and adaptive advantages.
π Abstract
Clustering is a fundamental problem in statistics, with applications across many scientific disciplines. In many modern applications involving clustering, the primary dataset (the target data) is accompanied by related datasets (the source data). Transferring information from such sources may improve clustering accuracy in the target, making transfer learning for clustering practically important. Despite recent progress, the conditions under which source data improve target clustering remain unclear in high-dimensional settings, even for the canonical Gaussian mixture model. In this paper, we study the clustering problem in a two-community Gaussian mixture model where relatedness is captured by the geometric alignment of the target and source cluster means. We develop a minimax-optimal transfer-assisted clustering procedure and characterize, up to logarithmic factors, the phase transition for consistent target clustering in terms of the signal-to-noise ratios, sample sizes, ambient dimension, and degree of alignment between the datasets. The technique is also extended to adaptively choose between the target-only or the source assisted clustering depending on the target signal strength. Furthermore, we also extend our techniques to accommodate multiple communities and and multiple source datasets. Extensive simulations and an analysis of a human lung single-cell RNA-sequencing atlas demonstrate the practical effectiveness of our methods.