BlockingPy: approximate nearest neighbours for blocking of records for entity resolution

📅 2025-04-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In entity resolution, traditional deterministic blocking methods suffer from low recall and poor precision when unique identifiers are unavailable, primarily due to field errors or omissions. To address this, we propose a fault-tolerant blocking framework that integrates approximate nearest neighbor (ANN) search with graph-based algorithms. Specifically, we introduce the first approach that jointly leverages Locality-Sensitive Hashing (LSH) and Hierarchical Navigable Small World (HNSW) indexing within connected-component analysis to generate candidate pairs with high recall and low redundancy. The framework unifies support for diverse vector embeddings and similarity metrics via standardized interfaces. Evaluated on official benchmark datasets, our method achieves 12–28% higher recall compared to conventional blocking techniques while reducing the number of pairwise comparisons by over 90%, thereby significantly improving both efficiency and accuracy.

Technology Category

Search and Optimization: Distributed SearchData Mining & Knowledge Management: Linked Open Data, Knowledge Graphs & KB CompletionNatural Language Processing: Information Extraction

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Entity resolution (probabilistic record linkage, deduplication) is a key step in scientific analysis and data science pipelines involving multiple data sources. The objective of entity resolution is to link records without common unique identifiers that refer to the same entity (e.g., person, company). However, without identifiers, researchers need to specify which records to compare in order to calculate matching probability and reduce computational complexity. One solution is to deterministically block records based on some common variables, such as names, dates of birth or sex or use phonetic algorithms. However, this approach assumes that these variables are free of errors and completely observed, which is often not the case. To address this challenge, we have developed a Python package, BlockingPy, which uses blocking via modern approximate nearest neighbour search and graph algorithms to reduce the number of comparisons. In this paper, we present the design of the package, its functionalities and two case studies related to official statistics. The presented software will be useful for researchers interested in linking data from various sources.
Problem

Research questions and friction points this paper is trying to address.

Linking records without common unique identifiers
Reducing computational complexity in entity resolution
Handling errors and missing data in blocking variables
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses approximate nearest neighbour search
Employs graph algorithms for blocking
Reduces comparisons in entity resolution
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Tymoteusz Strojny
Institute of Informatics and Electronic Economy, Poznań University of Economics and Business, Al. Niepodległości 10, 61-875 Poznań, Poland; Centre for the Methodology of Population Studies, Statistical Office in Poznań, Wojska Polskiego 27/29, 60-624 Poznań, Poland
M
Maciej Berkesewicz
Department of Statistics, Poznań University of Economics and Business, Al. Niepodległości 10, 61-875 Poznań, Poland; Centre for the Methodology of Population Studies, Statistical Office in Poznań, Wojska Polskiego 27/29, 60-624 Poznań, Poland