🤖 AI Summary
This study addresses the challenges of likelihood construction and the absence of spatial contiguity constraints in clustering complex spatial objects by proposing a Bayesian clustering framework based on distance matrices and spatial adjacency graphs. The method achieves spatially contiguous clustering through a hierarchical distance likelihood coupled with a random graph partition prior, while introducing node fragility parameters to characterize intra-cluster dependencies and provide interpretable measures of object centrality. Posterior inference is efficiently performed via a partially collapsed Markov chain Monte Carlo algorithm. Simulation experiments demonstrate that the proposed approach significantly improves region recovery accuracy, and its practical utility is validated through applications to population distribution and cancer mortality data.
📝 Abstract
Clustering problems increasingly involve complex objects observed over space, such as distributions, matrices, functions, images, or multivariate data, for which a scientifically meaningful dissimilarity between objects is often easier to specify and computationally more tractable than an object-response likelihood model. We propose DISCCO, a Bayesian framework for clustering spatially indexed complex objects using only a pairwise distance matrix and a spatial adjacency graph, with broad applicability and minimal user modeling requirements. The model combines a hierarchical distance-based likelihood accounting for within-cluster compactness and between-cluster separation with a random spatial graph partition prior, ensuring that posterior clusters are spatially contiguous. A key model feature is a set of node-specific frailty parameters that induce dependence among overlapping within-cluster distances and provide posterior summaries of object-level centrality or peripherality within each inferred cluster. We develop a partially collapsed Markov chain Monte Carlo algorithm for posterior inference. Simulations with distribution- and matrix-valued responses show that the proposed spatial distance-clustering framework improves region recovery relative to existing distance-clustering methods, while the frailty layer provides interpretable centrality summaries. Real applications to Houston Census Block Group racial-composition distributions and Western US county-level cancer mortality matrices illustrate how the method recovers interpretable contiguous clusters and frailty-based centrality maps.