🤖 AI Summary
This work addresses the limitations of existing spatial transcriptomics clustering methods, which typically decouple dimensionality reduction from clustering and operate at a single scale, thereby struggling to simultaneously resolve multiscale structures such as cell types and tissue regions. To overcome this, the authors propose BayesClint—a Bayesian multiscale clustering framework that uniquely integrates factor analysis, multi-sample alignment, spatially constrained clustering, and posterior-probability–based feature selection into a unified model. This approach enables simultaneous inference of both cellular- and tissue-level structures. By leveraging sparse Bayesian estimation of factor loadings, BayesClint effectively identifies active and differentially expressed genes. Comprehensive evaluations on both simulated and real datasets demonstrate that BayesClint significantly outperforms current state-of-the-art methods, achieving higher clustering accuracy and enhanced biological interpretability.
📝 Abstract
Recent advances in spatial transcriptomics have enabled researchers to profile gene expression at the single-cell spatial resolution, often for multiple tissue samples in a single study. This high-dimensional molecular profile for each cell can be used to sort cells into cell types with distinct functions, or segment the tissue into biologically relevant spatial domains. Although many non-spatial and spatial clustering methods have been developed to cluster these cells into cell types or spatial domains, most have two main limitations: first, they perform dimension reduction and clustering separately; second, they cluster cells at a single scale, rather than treating cell type and spatial domain clustering as distinct tasks at two different scales. To overcome these limitations, we propose BayesClint, a Bayesian method that simultaneously performs factor analysis and spatial clustering on multiple samples, where the clustering is done jointly at the single-cell and tissue regional scale. To increase interpretability, we employ a feature selection mechanism within the estimation of the sparse factor loadings matrix, which detects active genes and differentially expressed genes that discriminate between cell type clusters. We illustrate the advantages of the method over alternative state-of-the-art approaches through simulation studies and two real data applications.