adaptive spatial blocking

Designs and implements methods that partition spatial data into adaptive, disjoint local blocks and construct block-level summaries and hypothesis tests for spatial structure or clustering. This includes algorithms to choose block sizes and shapes under point-count/shape constraints, aggregate blockwise evidence into global inferences, and avoid expensive all-pairs spatial computations (e.g., replacing Ripley K style operations).

adaptivespatialblocking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the high computational cost of large-scale spatial point pattern clustering inference, which hinders its application to high-throughput spatial proteomics data. The authors propose an efficient testing framework based on adaptive spatial tiling: by extracting non-overlapping local tiles that satisfy constraints on both point count and geometric shape, and combining Ripley’s K-function with asymptotic normal approximation and evidence aggregation across multiple tiles, the method enables scalable clustering inference and rapid p-value computation. While preserving statistical power, the approach achieves substantial gains in computational efficiency. It successfully detects spatial clustering of plasma cells and their co-localization with macrophages in both simulated data and real human gut spatial proteomics datasets, demonstrating strong scalability and practical utility.

computational efficiencyhigh-throughput spatial proteomicsRipley's K-function

Geographic hotspot prediction suffers from weak high-dimensional spatial representation and low pattern recognition accuracy in existing methods. To address this, we propose a novel point cloud–voxel–community joint clustering paradigm: geographic events are modeled as 3D point clouds; spatial discretization is achieved via voxelization; a spatial similarity graph is constructed; and graph neural network–driven community detection is introduced to uncover multi-scale topological structural features. This approach pioneers point-cloud-driven voxel-level community partitioning, synergistically integrating geometric representation with topological analysis—thereby overcoming the modeling limitations of conventional grid- or density-based clustering for complex spatial patterns. Evaluated on the Turkey archaeological site dataset, our method achieves a 19.31% speedup over K-means and DBSCAN baselines, with only a 6% accuracy degradation, demonstrating superior efficiency and robustness.

Enhancing hotspot detection accuracy with community partitioningImproving analytical efficiency in high-dimensional spatial dataPredicting geographical hotspots using point cloud-voxel clustering

Informed Random Partition Models with Temporal Dependence

Nov 24, 2023
SP
Sally Paganin
🏛️ The Ohio State University | Brigham Young University | Pontificia Universidad Católica de Chile

In model-based clustering, incomplete prior knowledge and difficulty in specifying appropriate uncertainty over the entire partition hinder robust inference. Method: We propose a locally weighted probabilistic modeling framework that allocates prior uncertainty at the individual observation level—departing from conventional global penalty schemes—and integrates spatiotemporal dependence via a spatiotemporal Gaussian process, coupled with a Bayesian nonparametric random partition model and MCMC inference. Contribution/Results: Our approach enables fine-grained, subset-specific uncertainty encoding, substantially enhancing flexibility and robustness in incorporating expert knowledge. Evaluated on PM₁₀ spatiotemporal data and synthetic experiments, it achieves 12–19% higher clustering accuracy than baseline methods and demonstrates superior robustness to misspecified priors.

Addresses spatio-temporal data clustering with adaptive probability weightsEnhances clustering with flexible expert-informed uncertainty levelsImproves prior information handling in partition-based clustering

Clustering with minimum spanning trees: How good can it be?

Mar 10, 2023
MG
M. Gagolewski
🏛️ Polish Academy of Sciences | Warsaw University of Technology | QED Software

Minimum spanning tree (MST)-based clustering lacks theoretical consistency guarantees and robust partitioning mechanisms, especially in low-dimensional divisive clustering. Method: We propose an enhanced Genie–information-theoretic MST clustering framework integrating dynamic edge pruning, entropy-guided hierarchical splitting, and MST post-processing techniques; we further design an Oracle consistency metric to quantify clustering fidelity against ground-truth partitions. Contribution/Results: Evaluated across multiple benchmark datasets, our method significantly outperforms mainstream non-MST algorithms—including K-means, spectral clustering, and DBSCAN—and approaches the Oracle performance upper bound—defined by expert-annotated ground truth—in several scenarios. Results demonstrate that the MST paradigm achieves strong competitiveness and scalability, offering both theoretical insight and practical utility for graph-structured unsupervised learning.

Comparing MST methods against expert labels and benchmarksDeveloping improved MST-based partitioning schemes for clusteringQuantifying MST effectiveness in low-dimensional clustering tasks

Automatic Parameter Selection for Non-Redundant Clustering

Dec 19, 2023
CL
Collin Leiber
🏛️ LMU Munich | University of Vienna

Subspace clustering in high-dimensional data often yields multiple semantically distinct subspaces, yet existing methods require manual specification of both the number of subspaces and the number of clusters within each—rendering them parameter-sensitive and poorly interpretable. This paper proposes an automatic, non-redundant multi-subspace clustering framework. First, it introduces the Minimum Description Length (MDL) principle to non-redundant clustering, enabling joint, adaptive inference of both the optimal number of subspaces and the cluster count per subspace. Second, it designs a split-merge-based greedy search strategy coupled with a subspace-level outlier encoding mechanism, allowing simultaneous outlier detection. Evaluated on multiple benchmark datasets, the method achieves competitive accuracy against state-of-the-art approaches while significantly improving parameter robustness, model interpretability, and practical applicability.

Automatically selects parameters for non-redundant clusteringDetects subspaces and clusters without user inputIdentifies outliers within each subspace efficiently

Latest Papers

What's happening recently
View more

This work addresses the limitations of traditional clustering algorithms—specifically, the high computational complexity (O(N²)) of density-based methods like DBSCAN and the inability of partitioning approaches such as K-Means to capture nonlinear structures or handle noise effectively. To overcome these challenges, the authors propose K-SCAN, a novel algorithm that synergistically integrates vector quantization with density-based analysis. K-SCAN first employs stochastic mini-batch K-Means to generate weighted micro-clusters and then performs density connectivity analysis on these micro-clusters. This approach achieves linear time complexity while accurately identifying nonlinear manifold structures and exhibiting robustness to noise. Experimental results demonstrate that K-SCAN runs over three times faster than BIRCH on million-scale datasets, attains an Adjusted Rand Index exceeding 0.99, and remains effective even with noise levels as high as 55%.

density-based clusteringnoise robustnessnon-linear clusters

This study addresses the mismatch between areal-aggregated geographic data and point-based spatial scan statistics, where representing regions by their centroids discards critical spatial information and reduces statistical power. To mitigate this limitation, the authors propose a simple yet scalable preprocessing strategy: uniformly sampling 20–50 points within each region’s geometry and distributing the region’s observed count equally among these points. This approach better preserves the underlying spatial distribution while remaining computationally tractable. Empirical evaluations demonstrate that the method substantially enhances the detection performance of spatial scan statistics on aggregated regional data across diverse scenarios. The authors advocate its adoption as a standard preprocessing step for analyzing areal-aggregated datasets in spatial anomaly detection tasks.

anomaly detectiongeospatial dataregion-aggregated data

This study addresses the problem of partitioning a polygon into the minimum number of strips of width at most 1, aligned with a given orthogonal direction, and producing a compact representation of the optimal partition. It introduces, for the first time in this domain, the Clarke–Cormack–Burkowski lattice-theoretic framework, modeling the problem via interval antichains and combining meet/join operations with dynamic programming to devise an input-sensitive optimal algorithm. For convex polygons, the approach achieves an O(log n)-time decision version and an O(h log(1 + n/h))-time reporting version, where h is the number of strips in the optimal solution. For both simple and self-overlapping polygons, it attains O(n log n) time complexity, while establishing matching lower bounds of Ω(n) and Ω(n log n), respectively, thereby yielding tight complexity characterizations for all three polygon classes.

computational geometrylattice theorylower bounds

This study addresses the problem of testing statistical hypotheses concerning the independence or homogeneity between marks and spatial locations in marked point processes at local scales. It proposes a chi-square-type test statistic based on a local inhomogeneous mark-weighted K-function—a novel application of such K-functions for constructing test statistics. The proposed method simultaneously detects both global and local departures from the null hypothesis and maintains high sensitivity even when mark structures are weak or sample sizes are limited. Empirical validation using real-world datasets, including forest ecology and seismic event data, demonstrates its effectiveness in identifying spatially dependent mark structures in complex scenarios, substantially improving the accuracy and applicability of local pattern inference.

hypothesis testinginhomogeneitylocal structure

Traditional marked point process analyses rely on global statistics and struggle to capture local spatial heterogeneity. This study proposes the Local Indicator of Mark Association (LIMA) framework, which for the first time incorporates compositional marks into local spatial analysis. By leveraging the centered log-ratio (clr) transformation and Aitchison geometry, LIMA maps compositional data into Euclidean space, enabling a pointwise decomposition of mark structure. The method effectively uncovers local clustering and “siphoning” effects that are obscured by global approaches. Simulation experiments demonstrate that LIMA substantially outperforms existing global methods in detecting local clusters. When applied to economic data from Castilla–La Mancha, Spain, LIMA successfully reveals latent regional economic agglomeration patterns.

Aitchison geometrycomposition-valued markslocal heterogeneity

Hot Scholars

GP

George Papadakis

National and Kapodistrian University of Athens
Entity ResolutionWeb Data MiningData Management
CW

Chang Wen Chen

Chair Professor of Visual Computing, The Hong Kong Polytechnic University
multimedia communicationmultimedia systemsimage/video processingmultimedia signal processing
BB

Burcu B. Keskin

University of Alabama
network designinventory managementresilient networksillicit supply chains
AC

Amelie Chi Zhou

Assistant Professor, HKBU, Hong Kong
High performance computingCloud computingBig data analytics
GJ

Gregory J. Bott

Associate Professor, University of Alabama
Human TraffickingInformation SecurityInformation Privacy