Selective Inference for Deep Clustering in Latent Spaces

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of assessing statistical reliability in deep clustering, where data selection bias and the limitations of conventional direct feature clustering methods hinder valid inference. Focusing on latent space clustering with pretrained encoders, this work introduces the first selective inference framework tailored to this setting. By correcting for complex selection events induced by nonlinear transformations, the proposed approach enables statistically valid hypothesis testing for cluster differences. Empirically, the method rigorously controls the Type I error rate while achieving substantially higher statistical power compared to conservative baselines. Furthermore, its application in genomics successfully identifies significant inter-cluster differences, demonstrating both theoretical soundness and practical utility for reliable post-clustering analysis in high-dimensional representation spaces.
📝 Abstract
Deep clustering is a powerful approach for discovering meaningful structures in high-dimensional data by learning a low-dimensional latent representation prior to clustering. Despite its empirical success, assessing the statistical reliability of the resulting clusters remains challenging. Testing discovered clusters on the same data induces selection bias and invalidates classical $p$-values. Selective inference (SI) provides a principled framework for correcting this bias, but existing methods focus on clustering performed directly on the observed features. In this work, we develop an SI framework for deep clustering with a fixed pretrained encoder. The key challenge is that cluster assignments are determined through a nonlinear transformation from the original data space to the latent space, resulting in a substantially more complex selection process than in conventional clustering. Our method provides a computationally tractable way to account for this process and enables valid statistical testing of differences between clusters identified in the latent space. Synthetic experiments demonstrate that the proposed method controls the Type I error rate while achieving higher power than valid but conservative baselines, and genomic applications show that it can identify significant cluster differences while appropriately accounting for selection bias. Our framework provides a principled approach to quantifying the statistical reliability of structures discovered by deep clustering.
Problem

Research questions and friction points this paper is trying to address.

Deep Clustering
Selective Inference
Latent Space
Selection Bias
Statistical Reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selective Inference
Deep Clustering
Latent Space
Nonlinear Transformation
Statistical Reliability
🔎 Similar Papers
No similar papers found.
E
Eina Mizui
Nagoya University
T
Tomohiro Shiraishi
Nagoya University, RIKEN
S
Shunichi Nishino
Nagoya University, RIKEN
Ichiro Takeuchi
Ichiro Takeuchi
Professor, Nagoya University
Machine LearningData ScienceArtificial Intelligence