A flexible model for record linkage

📅 2024-07-09
🏛️ Journal of the Royal Statistical Society Series C: Applied Statistics
📈 Citations: 1
Influential: 0
📄 PDF

career value

242K/year
🤖 AI Summary
This study addresses the entity linking problem in multi-source heterogeneous data lacking unique identifiers. We propose a probabilistic record linkage method that balances accuracy and scalability. Methodologically, we introduce the Stochastic EM algorithm into latent-variable generative models for the first time, explicitly modeling dependencies among link decisions and enforcing one-to-one constraints, while enabling robust linking under variable-quality fields. Our approach innovatively supports dynamic precision–efficiency trade-offs, effectively handling real-world challenges such as information evolution, data entry errors, and low-quality attributes. Extensive evaluation on large-scale real-world healthcare data demonstrates high linkage accuracy; simulation experiments confirm strong robustness to noise and missing values. The open-source R package FlexRL has been released and deployed in production environments.

Technology Category

Application Category

📝 Abstract
Combining data from various sources empowers researchers to explore innovative questions, for example those raised by conducting healthcare monitoring studies. However, the lack of a unique identifier often poses challenges. Record linkage procedures determine whether pairs of observations collected on different occasions belong to the same individual using partially identifying variables (e.g. birth year, postal code). Existing methodologies typically involve a compromise between computational efficiency and accuracy. Traditional approaches simplify this task by condensing information, yet they neglect dependencies among linkage decisions and disregard the one-to-one relationship required to establish coherent links. Modern approaches offer a comprehensive representation of the data generation process, at the expense of computational overhead and reduced flexibility. We propose a flexible method, that adapts to varying data complexities, addressing registration errors and accommodating changes of the identifying information over time. Our approach balances accuracy and scalability, estimating the linkage using a Stochastic Expectation Maximization algorithm on a latent variable model. We illustrate the ability of our methodology to connect observations using large real data applications and demonstrate the robustness of our model to the linking variables quality in a simulation study. The proposed algorithm FlexRL is implemented and available in an open source R package.
Problem

Research questions and friction points this paper is trying to address.

Linking records without unique identifiers across datasets
Balancing computational efficiency with linkage accuracy
Handling data errors and temporal changes in identifiers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flexible model adapting to data complexities
Stochastic EM algorithm for linkage estimation
Balances accuracy and computational scalability
🔎 Similar Papers
No similar papers found.
K
Kayané Robach
Department of Epidemiology and Data Science, Amsterdam UMC location Vrije Universiteit Amsterdam, De Boelelaan 1117, 1081 HV Amsterdam, The Netherlands
S
Stéphanie L van der Pas
Department of Epidemiology and Data Science, Amsterdam UMC location Vrije Universiteit Amsterdam, De Boelelaan 1117, 1081 HV Amsterdam, The Netherlands
M
M. A. van de Wiel
Department of Epidemiology and Data Science, Amsterdam UMC location Vrije Universiteit Amsterdam, De Boelelaan 1117, 1081 HV Amsterdam, The Netherlands
M
Michel H. Hof
Amsterdam Public Health, Methodology, The Netherlands