Efficient Identification of High Similarity Clusters in Polygon Datasets

📅 2025-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the prohibitively high computational cost of identifying highly similar polygon clusters in large-scale spatial datasets, this paper proposes an efficient framework integrating dynamic similarity indexing, supervised scheduling, and recall-aware constraints. It innovatively employs kernel density estimation (KDE) to adaptively determine similarity thresholds and introduces a supervised learning model to prioritize candidate clusters, significantly reducing the number of clusters requiring expensive geometric validation while preserving both precision and recall. The framework leverages Shapely 2.0 for robust polygon operations and Triton for GPU-accelerated geometric computation, enabling end-to-end optimization. Experiments on datasets containing up to ten million polygons demonstrate strong scalability: computational cost is reduced by 42%–68%, while clustering accuracy remains above 95%. This work establishes a new, efficient, and reliable paradigm for similarity mining in geospatial big data.

Technology Category

Data Mining & Knowledge Management: Mining of Spatial, Temporal or Spatio-Temporal DataKnowledge Representation and Reasoning: Geometric, Spatial, and Temporal ReasoningSearch and Optimization: Distributed Search

Application Category

Graph Algorithms and Modeling for the Web: Querying, indexing, and retrieval in Web-related graphsWeb Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web textSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search engines
📝 Abstract
Advancements in tools like Shapely 2.0 and Triton can significantly improve the efficiency of spatial similarity computations by enabling faster and more scalable geometric operations. However, for extremely large datasets, these optimizations may face challenges due to the sheer volume of computations required. To address this, we propose a framework that reduces the number of clusters requiring verification, thereby decreasing the computational load on these systems. The framework integrates dynamic similarity index thresholding, supervised scheduling, and recall-constrained optimization to efficiently identify clusters with the highest spatial similarity while meeting user-defined precision and recall requirements. By leveraging Kernel Density Estimation (KDE) to dynamically determine similarity thresholds and machine learning models to prioritize clusters, our approach achieves substantial reductions in computational cost without sacrificing accuracy. Experimental results demonstrate the scalability and effectiveness of the method, offering a practical solution for large-scale geospatial analysis.
Problem

Research questions and friction points this paper is trying to address.

Reducing computational load for large-scale polygon similarity clustering
Identifying high spatial similarity clusters with precision constraints
Optimizing verification processes through dynamic thresholding and scheduling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic similarity thresholds using Kernel Density Estimation
Machine learning models for cluster prioritization scheduling
Recall-constrained optimization meeting precision requirements
🔎 Similar Papers
2024-03-06IEEE Transactions on Neural Networks and Learning SystemsCitations: 0
J
John N. Daras
Columbia University in the city of New York, New York, USA