🤖 AI Summary
To address the prohibitively high computational cost of identifying highly similar polygon clusters in large-scale spatial datasets, this paper proposes an efficient framework integrating dynamic similarity indexing, supervised scheduling, and recall-aware constraints. It innovatively employs kernel density estimation (KDE) to adaptively determine similarity thresholds and introduces a supervised learning model to prioritize candidate clusters, significantly reducing the number of clusters requiring expensive geometric validation while preserving both precision and recall. The framework leverages Shapely 2.0 for robust polygon operations and Triton for GPU-accelerated geometric computation, enabling end-to-end optimization. Experiments on datasets containing up to ten million polygons demonstrate strong scalability: computational cost is reduced by 42%–68%, while clustering accuracy remains above 95%. This work establishes a new, efficient, and reliable paradigm for similarity mining in geospatial big data.
📝 Abstract
Advancements in tools like Shapely 2.0 and Triton can significantly improve the efficiency of spatial similarity computations by enabling faster and more scalable geometric operations. However, for extremely large datasets, these optimizations may face challenges due to the sheer volume of computations required. To address this, we propose a framework that reduces the number of clusters requiring verification, thereby decreasing the computational load on these systems. The framework integrates dynamic similarity index thresholding, supervised scheduling, and recall-constrained optimization to efficiently identify clusters with the highest spatial similarity while meeting user-defined precision and recall requirements. By leveraging Kernel Density Estimation (KDE) to dynamically determine similarity thresholds and machine learning models to prioritize clusters, our approach achieves substantial reductions in computational cost without sacrificing accuracy. Experimental results demonstrate the scalability and effectiveness of the method, offering a practical solution for large-scale geospatial analysis.