Causal Secondary Analysis of Linked Data in the Presence of Mismatch Error

📅 2025-12-16
📈 Citations: 0
Influential: 0
📄 PDF

career value

204K/year
🤖 AI Summary
This paper addresses estimation bias in the average treatment effect (ATE) arising from record linkage mismatches in secondary analysis of multi-source observational data. We propose a novel causal inference method that, for the first time, formalizes linkage mismatch as a missing-data mechanism. Our approach constructs an estimating equation framework based on a two-component mixture model and an enhanced EM algorithm, enabling asymptotically consistent ATE estimation and valid statistical inference—even when linkage quality information is unavailable. Simulation studies and empirical applications demonstrate that the method substantially reduces estimation bias, improves ATE consistency, and achieves nominal coverage rates for confidence intervals—outperforming naive analyses that ignore linkage errors. By rigorously accounting for linkage uncertainty, our framework provides a generalizable, statistically principled solution for causal secondary analysis under imperfect record linkage.

Technology Category

Application Category

📝 Abstract
The increased prevalence of observational data and the need to integrate information from multiple sources are critical challenges in contemporary data analysis. Record linkage is a widely used tool for combining datasets in the absence of unique identifiers. The presence of linkage errors such as mismatched records, however, often hampers the analysis of data sets obtained in this way. This issue is more difficult to address in secondary analysis settings, where linkage and subsequent analysis are performed separately, and analysts have limited information about linkage quality. In this paper, we investigate the estimation of average treatment effects in the conventional potential outcome-based causal inference framework under linkage uncertainty. To mitigate the bias that would be incurred with naive analyses, we propose an approach based on estimating equations that treats the unknown match status indicators as missing data. Leveraging a variant of the Expectation-Maximization algorithm, these indicators are imputed based on a corresponding two-component mixture model. The approach is amenable to asymptotic inference. Simulation studies and a case study highlight the importance of accounting for linkage uncertainty and demonstrate the effectiveness of the proposed approach.
Problem

Research questions and friction points this paper is trying to address.

Estimates causal effects despite linkage errors in data integration.
Addresses bias from mismatched records in secondary data analysis.
Proposes a method to handle unknown match status in datasets.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Estimating equations treat match status as missing data
Expectation-Maximization variant imputes unknown match indicators
Two-component mixture model addresses linkage uncertainty bias
🔎 Similar Papers
No similar papers found.