🤖 AI Summary
This work proposes an interactive, human-in-the-loop visual clustering framework for high-dimensional data, addressing the limitations of static dimensionality reduction methods that often lack interpretability and cannot incorporate human prior knowledge. By introducing a closed-loop feedback mechanism, the approach enables users to dynamically guide nonlinear projections through a small number of must-link and cannot-link constraints, while simultaneously refining low-dimensional embeddings via semi-supervised clustering. The framework further supports traceability from clustering results back to the original feature space, facilitating interpretable analysis. Experimental results on multiple benchmark datasets demonstrate that just a few rounds of user interaction significantly improve clustering quality, achieving both efficiency and interpretability in high-dimensional clustering tasks.
📝 Abstract
High-dimensional datasets are increasingly common across scientific and industrial domains, yet they remain difficult to cluster effectively due to the diminishing usefulness of distance metrics and the tendency of clusters to collapse or overlap when projected into lower dimensions. Traditional dimensionality reduction techniques generate static 2D or 3D embeddings that provide limited interpretability and do not offer a mechanism to leverage the analyst's intuition during exploration. To address this gap, we propose Interactive Project-Based Clustering (IPBC), a framework that reframes clustering as an iterative human-guided visual analysis process. IPBC integrates a nonlinear projection module with a feedback loop that allows users to modify the embedding by adjusting viewing angles and supplying simple constraints such as must-link or cannot-link relationships. These constraints reshape the objective of the projection model, gradually pulling semantically related points closer together and pushing unrelated points further apart. As the projection becomes more structured and expressive through user interaction, a conventional clustering algorithm operating on the optimized 2D layout can more reliably identify distinct groups. An additional explainability component then maps each discovered cluster back to the original feature space, producing interpretable rules or feature rankings that highlight what distinguishes each cluster. Experiments on various benchmark datasets show that only a small number of interactive refinement steps can substantially improve cluster quality. Overall, IPBC turns clustering into a collaborative discovery process in which machine representation and human insight reinforce one another.