🤖 AI Summary
In entity resolution, traditional deterministic blocking methods suffer from low recall and poor precision when unique identifiers are unavailable, primarily due to field errors or omissions. To address this, we propose a fault-tolerant blocking framework that integrates approximate nearest neighbor (ANN) search with graph-based algorithms. Specifically, we introduce the first approach that jointly leverages Locality-Sensitive Hashing (LSH) and Hierarchical Navigable Small World (HNSW) indexing within connected-component analysis to generate candidate pairs with high recall and low redundancy. The framework unifies support for diverse vector embeddings and similarity metrics via standardized interfaces. Evaluated on official benchmark datasets, our method achieves 12–28% higher recall compared to conventional blocking techniques while reducing the number of pairwise comparisons by over 90%, thereby significantly improving both efficiency and accuracy.
📝 Abstract
Entity resolution (probabilistic record linkage, deduplication) is a key step in scientific analysis and data science pipelines involving multiple data sources. The objective of entity resolution is to link records without common unique identifiers that refer to the same entity (e.g., person, company). However, without identifiers, researchers need to specify which records to compare in order to calculate matching probability and reduce computational complexity. One solution is to deterministically block records based on some common variables, such as names, dates of birth or sex or use phonetic algorithms. However, this approach assumes that these variables are free of errors and completely observed, which is often not the case. To address this challenge, we have developed a Python package, BlockingPy, which uses blocking via modern approximate nearest neighbour search and graph algorithms to reduce the number of comparisons. In this paper, we present the design of the package, its functionalities and two case studies related to official statistics. The presented software will be useful for researchers interested in linking data from various sources.