Score
Designs, implements, and evaluates methods that group geographic units (points, polygons, or raster cells) into clusters or typologies based on spatial proximity and attribute similarity, producing maps and partitionings that reveal spatial patterns. Performs cluster validation and spatial analysis (e.g., assessing spatial autocorrelation, stability, and cluster boundaries) and interprets resulting typologies to inform location‑specific decisions or targeted actions.
To address the prohibitively high computational cost of identifying highly similar polygon clusters in large-scale spatial datasets, this paper proposes an efficient framework integrating dynamic similarity indexing, supervised scheduling, and recall-aware constraints. It innovatively employs kernel density estimation (KDE) to adaptively determine similarity thresholds and introduces a supervised learning model to prioritize candidate clusters, significantly reducing the number of clusters requiring expensive geometric validation while preserving both precision and recall. The framework leverages Shapely 2.0 for robust polygon operations and Triton for GPU-accelerated geometric computation, enabling end-to-end optimization. Experiments on datasets containing up to ten million polygons demonstrate strong scalability: computational cost is reduced by 42%–68%, while clustering accuracy remains above 95%. This work establishes a new, efficient, and reliable paradigm for similarity mining in geospatial big data.
Existing methods struggle to effectively identify spatial clusters in categorical functional data. This work proposes a novel spatial scan statistic that, for the first time, integrates an encoding mechanism for categorical functional data with a nonparametric scan statistic to detect clustering patterns in complex spatiotemporal datasets. The method demonstrates high true positive rates, low false positive rates, and high positive predictive values in simulation studies. It was successfully applied to winter 2024 air pollution data from France, effectively uncovering the spatial clustering structure of pollution events. This approach establishes a new paradigm for spatial analysis of categorical functional data.
ROC curves are frequently misapplied in spatial presence–absence or presence-only modeling to assess goodness-of-fit or predictive accuracy, despite being intrinsically limited to evaluating a model’s ability to rank point events and insensitive to model specification. Method: This paper clarifies the statistical foundations of ROC analysis within spatial point pattern analysis and establishes theoretical links to spatial hypothesis testing and model diagnostics. Building on this framework, we propose five novel ROC-based methodologies: variable selection, model comparison, point-type isolation analysis, baseline correction, and spatial case–control analysis. Implemented within the spatstat package by integrating spatial point process statistics with nonparametric inference, these methods are validated across multiple real-world spatial datasets. Contribution/Results: The proposed approaches extend the theoretical interpretation and practical applicability of ROC analysis in spatial statistics; associated code is integrated into the spatstat development version and scheduled for public release.
This paper identifies and systematically analyzes common challenges in spatial data science across mainstream programming languages—R, Python, and Julia—including inconsistent spherical geometry modeling, ambiguous spatial/temporal semantics, conflation of intensive and extensive attributes, poor interoperability between data cube and vector formats, complex cross-package dependencies, and a persistent divide between GIS and physical modeling communities. Through multi-language ecosystem surveys, cross-community comparative analysis, and software engineering abstraction, we propose, for the first time, a cross-language semantic framework for spatial operations. The framework formally defines support types (point vs. block), specifies attribute-type constraints on operation validity, and refactors spherical Simple Features logic. We distill five foundational insights that establish a methodological basis and practical guidance for tool interoperability, pedagogical alignment, and open-source governance in spatial computing.
High-stakes domains—such as healthcare and finance—demand interpretable clustering outcomes to ensure transparency, accountability, and regulatory compliance. Method: This survey systematically analyzes over 120 scholarly works, proposing the first unified taxonomy of interpretability dimensions for clustering. It rigorously distinguishes intrinsically interpretable models—including rule-based, prototype-based, and sparsity-driven approaches—from post-hoc explanation techniques—such as visualization, feature attribution, and local surrogate modeling. The study further develops a use-case-oriented, structured classification framework and principled evaluation criteria. Contribution/Results: It introduces the first practical guideline for selecting appropriate interpretable clustering methods based on application requirements. The work bridges theoretical foundations with real-world deployment, providing both conceptual clarity and actionable insights to support the development and adoption of clustering algorithms that jointly optimize accuracy and interpretability—thereby advancing trustworthy AI in ethically and regulatorily sensitive contexts.
This study addresses the challenge of accurately estimating the number of clusters in datasets with unknown clustering structure by proposing a nonparametric method based on pairwise distances. The approach generates p-values through dependence-adjusted multiple hypothesis testing that adapts to sample size and employs a stepwise selection strategy to infer the optimal number of clusters, without requiring a pre-specified cluster count. It is applicable to data of arbitrary dimensionality and compatible with existing clustering algorithms. Experimental results demonstrate that the method substantially reduces computational overhead while achieving superior accuracy and robustness in cluster number estimation compared to current state-of-the-art metrics, offering strong theoretical grounding and practical utility.
This study addresses the mismatch between areal-aggregated geographic data and point-based spatial scan statistics, where representing regions by their centroids discards critical spatial information and reduces statistical power. To mitigate this limitation, the authors propose a simple yet scalable preprocessing strategy: uniformly sampling 20–50 points within each region’s geometry and distributing the region’s observed count equally among these points. This approach better preserves the underlying spatial distribution while remaining computationally tractable. Empirical evaluations demonstrate that the method substantially enhances the detection performance of spatial scan statistics on aggregated regional data across diverse scenarios. The authors advocate its adoption as a standard preprocessing step for analyzing areal-aggregated datasets in spatial anomaly detection tasks.
Clustering analyses often lack quantitative assessment of reproducibility. To address this gap, this work proposes ERICA, the first systematic framework for quantifying clustering reproducibility. ERICA generates stability statistics through iterative cluster assignments and integrates quantitative visualization to reveal inter-cluster similarity and potential outliers. The method is validated on synthetic datasets and applied to breast cancer gene expression data, where it identifies subsets of clustering results that are irreproducible. These findings underscore ERICA’s critical value in real-world applications for evaluating the reliability of clustering outcomes and the robustness of underlying data structures.
This study addresses a critical gap in clustering interpretability: existing post-hoc explanation methods primarily focus on feature importance or instance-level explanations and struggle to reliably uncover structured patterns within clusters. To systematically evaluate this limitation, the authors conduct the first controlled assessment of multiple explanation techniques—including random forest permutation importance, LIME, and principal component analysis—in synthetic datasets where ground-truth structured patterns are explicitly embedded. Results demonstrate that while these methods partially recover relevant features, none consistently identifies all types of predefined patterns. This reveals a fundamental shortcoming of current interpretability tools in capturing pattern-level cluster structure and underscores the urgent need for dedicated methods designed specifically for detecting and explaining such intra-cluster patterns.
This study addresses the reliance on expert-derived weights and the absence of data-driven approaches in multi-thematic geographic layer fusion by proposing GIS-moGA, a novel bi-objective evolutionary framework. For the first time, spatial autocorrelation structure is integrated into a multi-objective genetic algorithm to automatically optimize layer weights by simultaneously maximizing global spatial autocorrelation (Global Moran’s I) and minimizing local spatial heterogeneity (LISA variance). The method employs a queen-contiguity sparse matrix to enhance computational efficiency for large-scale geographic units and uncovers the critical role of high mutation rates in maintaining population diversity. Validated on epidemiological data from 523 spatial units in Araraquara, Brazil, GIS-moGA significantly outperforms the Analytic Hierarchy Process (AHP), yielding a substantially larger Pareto front hypervolume and markedly improved spatial consistency (p < 0.001, Cliff’s delta = 0.87).