record linkage

Design and implement methods and pipelines to identify and merge records that represent the same real-world entities across two or more datasets, including blocking/candidate retrieval, construction of matching keys, similarity-scoring functions, multi-signal and probabilistic weighting schemes, and computation of composite match scores with confidence tiers. Validate linkage quality through manual review and quantitative metrics and produce cleaned, documented linked datasets suitable for downstream analysis.

recordlinkage

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

A flexible model for record linkage

Jul 09, 2024
KR
Kayané Robach
🏛️ Amsterdam UMC | Amsterdam Public Health

This study addresses the entity linking problem in multi-source heterogeneous data lacking unique identifiers. We propose a probabilistic record linkage method that balances accuracy and scalability. Methodologically, we introduce the Stochastic EM algorithm into latent-variable generative models for the first time, explicitly modeling dependencies among link decisions and enforcing one-to-one constraints, while enabling robust linking under variable-quality fields. Our approach innovatively supports dynamic precision–efficiency trade-offs, effectively handling real-world challenges such as information evolution, data entry errors, and low-quality attributes. Extensive evaluation on large-scale real-world healthcare data demonstrates high linkage accuracy; simulation experiments confirm strong robustness to noise and missing values. The open-source R package FlexRL has been released and deployed in production environments.

Balancing computational efficiency with linkage accuracyHandling data errors and temporal changes in identifiersLinking records without unique identifiers across datasets

This work addresses the poor generalizability of existing self-service entity resolution methods on unseen datasets and their difficulty in balancing precision and recall, which often leads to error propagation and cascading erroneous merges. To overcome these limitations, the authors propose a robust self-service entity resolution system that integrates three key innovations: automated selection among multiple algorithms, decoupled optimization of precision and recall—enforcing precision through rule-based veto mechanisms while enhancing recall via diverse candidate generation—and non-transitive cross-cluster merge validation to prevent error propagation. Extensive experiments on six real-world benchmark datasets, ranging from 864 to 5 million records, demonstrate that the proposed approach significantly reduces incorrect merges, enhances practical utility, and substantially lowers the trial-and-error burden for practitioners.

Entity ResolutionFalse-Positive LinkMatching Algorithm

BlockingPy: approximate nearest neighbours for blocking of records for entity resolution

Apr 05, 2025
TS
Tymoteusz Strojny
🏛️ Poznań University of Economics and Business | Statistical Office in Poznań

In entity resolution, traditional deterministic blocking methods suffer from low recall and poor precision when unique identifiers are unavailable, primarily due to field errors or omissions. To address this, we propose a fault-tolerant blocking framework that integrates approximate nearest neighbor (ANN) search with graph-based algorithms. Specifically, we introduce the first approach that jointly leverages Locality-Sensitive Hashing (LSH) and Hierarchical Navigable Small World (HNSW) indexing within connected-component analysis to generate candidate pairs with high recall and low redundancy. The framework unifies support for diverse vector embeddings and similarity metrics via standardized interfaces. Evaluated on official benchmark datasets, our method achieves 12–28% higher recall compared to conventional blocking techniques while reducing the number of pairwise comparisons by over 90%, thereby significantly improving both efficiency and accuracy.

Handling errors and missing data in blocking variablesLinking records without common unique identifiersReducing computational complexity in entity resolution

Contextual Graph Embeddings: Accounting for Data Characteristics in Heterogeneous Data Integration

Nov 12, 2025
YH
Yuka Haruki
🏛️ The University of Tokyo | Infomart Corporation

In heterogeneous data integration, schema matching and entity resolution are significantly affected by domain-specific characteristics, data scale, missingness rates, and attribute overlap—yet existing graph-based methods struggle to jointly leverage structural and semantic information. Method: This paper proposes a context-aware graph embedding framework that unifies tabular structure, column-level textual descriptions, and external knowledge, employing graph neural networks for joint encoding and embedding learning of multi-source heterogeneous data. Contribution/Results: A key innovation is the context-enhancement mechanism, which systematically uncovers how data characteristics influence matching performance and empirically demonstrates that contextual modeling substantially improves robustness and accuracy—especially under challenging conditions such as high missingness rates and prevalent numeric columns. Extensive experiments across multiple domain-specific benchmark datasets show that our method consistently outperforms state-of-the-art graph-based baselines.

Addressing dataset characteristics' impact on data integration effectivenessAutomating schema matching and entity resolution in heterogeneous datasetsImproving matching reliability with contextual graph embeddings

Methods for quantifying dataset similarity: a review, taxonomy and comparison

Dec 07, 2023
MS
Marieke Stolte
🏛️ TU Dortmund University

This study addresses the critical need for quantifying dataset similarity in model generalization, transfer learning, simulation calibration, and two-sample testing. We systematically survey 118 similarity quantification methods and propose the first ten-dimensional classification framework, organizing approaches into seven technical categories: statistical distances (e.g., Wasserstein distance, Maximum Mean Discrepancy), kernel-based methods, information-theoretic measures, dimensionality-reduction embeddings, permutation tests, generative-model-based discriminators, and Gaussian process likelihood ratios. We develop a multi-dimensional evaluation system balancing theoretical guarantees, interpretability, and practical applicability, yielding a structured recommendation matrix aligned with task requirements and data characteristics. Furthermore, we introduce the first open-source, interactive tool for method selection—enabling real-time filtering and parameter configuration—to significantly enhance both selection efficiency and deployment suitability.

Compare 118 methods on applicability and interpretabilityProvide recommendations for selecting dataset similarity measuresReview and classify methods for quantifying dataset similarity

Latest Papers

What's happening recently
View more

This work addresses the joint schema-value matching problem between pivot tables and relational tables in data lakes, which demands semantic consistency, value compatibility, and generalization under anonymized data. To this end, we propose PiLLar, a novel framework that, for the first time, formulates the task as a large language model (LLM)-guided Monte Carlo Tree Search (MCTS), enabling unsupervised, training-free cross-domain adaptation. We provide a dynamic theoretical error analysis that guarantees asymptotic convergence. Furthermore, we construct PTbench, the first real-world benchmark for this problem. Experimental results show that PiLLar achieves an average matching accuracy of 87.94% on PTbench, significantly outperforming existing methods and demonstrating its effectiveness and strong generalization capability.

data anonymizationdata integrationpivot table schema matching

This work addresses the lack of end-to-end public benchmarks for data integration by introducing MaDI-Bench, the first comprehensive benchmark that spans the entire relational table integration pipeline—including schema matching, value normalization, entity blocking, entity matching, and data fusion. To mitigate benchmark saturation, MaDI-Bench incorporates a set of foundational cross-domain tasks along with a mechanism for generating extensible task variants. The benchmark supports both step-wise and end-to-end evaluation of system performance, validated through diverse pipelines ranging from manual and optimal combinations to large language model (LLM)-based approaches. All resources are publicly released to foster reproducible and holistic assessment of data integration systems.

data fusiondata integrationend-to-end benchmark

This study addresses the high cost and lack of systematic design in manual review of candidate record pairs, which hinders the simultaneous optimization of accuracy, representativeness, and uncertainty coverage under limited budgets. The authors model review as finite-population sampling based on fine-grained stratification and introduce a novel multidimensional stratification framework that integrates match-score intervals, comparison patterns, record-level ambiguity, and demographic groups. Ambiguity is quantified using match probability bands—derived from deciles of model scores—and conditional candidate perplexity. A tunable allocation mechanism is achieved through within-band error tolerances and global budget scaling. Experiments demonstrate that by reviewing only 7% of samples (compared to 23% in the baseline), the framework preserves high-segment matching accuracy and ambiguity distribution, confirming its effectiveness and flexibility under resource constraints.

ambiguity-awareclerical reviewdeduplication

This work addresses the high computational and communication overhead that hinders scalability in two-party private record linkage (PPRL) at large scale. To overcome this limitation, the authors propose a “filter-then-link” framework that introduces a lightweight privacy-preserving record screening (PPRS) phase prior to PPRL, enabling efficient assessment of collaboration value. The key innovation lies in an Oblivious Attribute/Feature Alignment protocol that supports approximate matching and pattern awareness, thereby transcending the symmetric-function constraint inherent in circuit-based private set intersection (PSI). Building upon circuit PSI, they further develop an Appraisal system that integrates secure shuffling and PSI techniques to achieve highly efficient screening. Experimental results demonstrate that, under identical constraints, their approach handles up to 850× more records and operates 165× faster than the state-of-the-art SFour system, substantially accelerating the identification of high-value collaboration partners.

Computational OverheadData CollaborationPrivacy-Preserving Record Linkage

Hot Scholars

PM

Philipp Mayr

GESIS - Leibniz Institute for the Social Sciences
Interactive Information RetrievalInformetricsDigital librariesInformation Seeking
YD

Yang Ding

Shanghai Jiao Tong university
3D ReconstructionMedical Image Process
SP

Silvio Peroni

University of Bologna
Semantic PublishingSemantic WebOpen ScienceScience of Science
FO

Francesco Osborne

KMi, The Open University
Science of ScienceInformation ExtractionKnowledge GraphsArtificial Intelligence
EM

Enrico Motta

Professor of Knowledge Technologies, KMi, The Open University
Semantic WebOntology EngineeringKnowledge SystemsData Science