Relaxing the Assumption of Strongly Non-Informative Linkage Error in Secondary Regression Analysis of Linked Files

📅 2025-10-20
📈 Citations: 0
Influential: 0
📄 PDF

career value

204K/year
🤖 AI Summary
Multi-source record linkage often introduces non-ignorable errors—such as missed links or false matches—that induce bias in secondary analyses. Conventional methods rely on the strong non-informative linkage assumption, i.e., that linkage errors are independent of analysis covariates—a condition frequently violated in practice. This paper proposes a two-component mixture model framework that relaxes this assumption by allowing the linkage mechanism to depend on observable covariates. The method enables error correction in secondary analyses where only linked data are available and no details about the linkage process are known. By jointly modeling linkage uncertainty and regression analysis, we derive consistent estimators and provide a practical EM algorithm for implementation. Simulation studies and empirical applications demonstrate that the proposed approach substantially improves parameter estimation accuracy and inferential robustness under non-ignorable linkage errors, thereby broadening the scope of valid analyses for linked data.

Technology Category

Application Category

📝 Abstract
Data analysis of files that are a result of linking records from multiple sources are often affected by linkage errors. Records may be linked incorrectly, or their links may be missed. In consequence, it is essential that such errors are taken into account to ensure valid post-linkage inference. Here, we propose an extension to a general framework for regression with linked covariates and responses based on a two-component mixture model, which was developed in prior work. This framework addresses the challenging case of secondary analysis in which only the linked data is available and information about the record linkage process is limited. The extension considered herein relaxes the assumption of strongly non-informative linkage in the framework according to which linkage does not depend on the covariates used in the analysis, which may be limiting in practice. The effectiveness of the proposed extension is investigated by simulations and a case study.
Problem

Research questions and friction points this paper is trying to address.

Addresses linkage errors in regression analysis of linked data files
Relaxes restrictive assumption of non-informative linkage errors
Enables valid inference when linkage process information is limited
Innovation

Methods, ideas, or system contributions that make the work stand out.

Extends mixture model framework for linked data regression
Relaxes strong non-informative linkage error assumption
Validates method through simulations and case study