🤖 AI Summary
This study addresses the challenge of assessing statistical reliability in deep clustering, where data selection bias and the limitations of conventional direct feature clustering methods hinder valid inference. Focusing on latent space clustering with pretrained encoders, this work introduces the first selective inference framework tailored to this setting. By correcting for complex selection events induced by nonlinear transformations, the proposed approach enables statistically valid hypothesis testing for cluster differences. Empirically, the method rigorously controls the Type I error rate while achieving substantially higher statistical power compared to conservative baselines. Furthermore, its application in genomics successfully identifies significant inter-cluster differences, demonstrating both theoretical soundness and practical utility for reliable post-clustering analysis in high-dimensional representation spaces.
📝 Abstract
Deep clustering is a powerful approach for discovering meaningful structures in high-dimensional data by learning a low-dimensional latent representation prior to clustering. Despite its empirical success, assessing the statistical reliability of the resulting clusters remains challenging. Testing discovered clusters on the same data induces selection bias and invalidates classical $p$-values. Selective inference (SI) provides a principled framework for correcting this bias, but existing methods focus on clustering performed directly on the observed features. In this work, we develop an SI framework for deep clustering with a fixed pretrained encoder. The key challenge is that cluster assignments are determined through a nonlinear transformation from the original data space to the latent space, resulting in a substantially more complex selection process than in conventional clustering. Our method provides a computationally tractable way to account for this process and enables valid statistical testing of differences between clusters identified in the latent space. Synthetic experiments demonstrate that the proposed method controls the Type I error rate while achieving higher power than valid but conservative baselines, and genomic applications show that it can identify significant cluster differences while appropriately accounting for selection bias. Our framework provides a principled approach to quantifying the statistical reliability of structures discovered by deep clustering.