administrative data linkage

Methods for linking, validating, harmonizing, and merging administrative records and survey data across sources and time to construct analysis-ready measures for causal and descriptive studies.

administrativedatalinkage

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Unfolding Data Quality Dimensions in Practice: A Survey

Jul 23, 2025
VP
Vasileios Papastergios
🏛️ Aristotle University | Hasso Plattner Institute | University of Potsdam

A significant gap exists between theoretical definitions of data quality dimensions—such as accuracy, completeness, consistency, and timeliness—as stipulated in standards (e.g., ISO/IEC 25012) and the actual functionalities implemented in widely used data quality tools, with no systematic mapping analysis to date. Method: We conducted a systematic literature review and performed functional reverse engineering on seven mainstream open-source data quality tools. Contribution/Results: We present the first many-to-many mapping framework linking data quality dimensions to concrete tool capabilities, introducing a cross-dimensional, fine-grained functional categorization schema and generating a structured correspondence matrix covering all core dimensions. This work bridges the theory–practice divide, delivers an actionable guide for data quality assessment, and substantially enhances the scientific rigor and reusability of tool selection, functionality design, and standard implementation.

Bridging gap between data quality theory and practiceMapping tool functionalities to data quality dimensionsProviding unified view on fragmented quality checks

Must-Read Papers

Most classic and influential ideas
View more

A flexible model for record linkage

Jul 09, 2024
KR
Kayané Robach
🏛️ Amsterdam UMC | Amsterdam Public Health

This study addresses the entity linking problem in multi-source heterogeneous data lacking unique identifiers. We propose a probabilistic record linkage method that balances accuracy and scalability. Methodologically, we introduce the Stochastic EM algorithm into latent-variable generative models for the first time, explicitly modeling dependencies among link decisions and enforcing one-to-one constraints, while enabling robust linking under variable-quality fields. Our approach innovatively supports dynamic precision–efficiency trade-offs, effectively handling real-world challenges such as information evolution, data entry errors, and low-quality attributes. Extensive evaluation on large-scale real-world healthcare data demonstrates high linkage accuracy; simulation experiments confirm strong robustness to noise and missing values. The open-source R package FlexRL has been released and deployed in production environments.

Balancing computational efficiency with linkage accuracyHandling data errors and temporal changes in identifiersLinking records without unique identifiers across datasets

Integrating smart surveys with traditional surveys

Oct 08, 2025
DM
Danielle Mccool
🏛️ Utrecht University

This study addresses the inherent differences in measurement mechanisms and sample representativeness between intelligent surveys (e.g., sensor- and AI-driven data collection) and traditional surveys. Methodologically, it distinguishes two integration paradigms: “hybrid-mode” integration—emphasizing outcome alignment to enable direct data merging—and “multi-source integration”—fusing heterogeneous data during modeling to leverage complementary strengths. A context-driven decision framework is proposed, integrating sensing technologies, machine learning, and statistical modeling for synergistic analysis. The key contribution lies in the first systematic delineation of the applicability boundaries and trade-offs between these two paradigms, offering actionable strategies for official statistics integration. Empirical application demonstrates substantial improvements in data quality and timeliness for dynamic behavioral surveys—particularly travel behavior—thereby advancing theoretical foundations and practical guidelines for statistical modernization in multi-source data environments. (149 words)

Addressing measurement differences between smart and traditional surveysDeveloping a framework for selecting integration strategiesIntegrating smart surveys with traditional diary surveys

To address the challenge of achieving FAIR interoperability and reusability for multi-source, heterogeneous scientific data—such as the RADx COVID-19 response data—in large-scale environments, this paper proposes a general-purpose, reproducible, and extensible data harmonization framework. Our approach introduces a novel harmonization paradigm based on parameterized primitive operations and automated execution tracing. It integrates a customizable data representation model, a configurable operation library, and mechanisms for transformation logging and dependency tracking—ensuring protocol reproducibility, process auditability, and transformation reusability. Evaluated in real-world deployment within the RADx Data Hub, the framework significantly lowers the barrier to entry for domain experts, improves harmonization efficiency and transparency, and enables high-quality cross-study analyses. This work provides a scalable, transferable technical pathway for FAIR-compliant data integration across diverse scientific domains.

Addresses challenges in aligning heterogeneous data setsFacilitates data harmonization for FAIR principles complianceSupports reproducible and scalable data integration processes

Combining Experimental and Observational Data to Estimate Treatment Effects on Long Term Outcomes

Jun 17, 2020
SA
S. Athey
🏛️ Stanford University | Harvard University | NBER

This study addresses selection bias in estimating long-term causal effects—such as graduation rates—from observational studies. We propose a novel control function approach that leverages experimental estimates of treatment effects on short-term outcomes (e.g., eighth-grade test scores) to correct for unobserved confounding in large-scale administrative observational data. Our method integrates insights from difference-in-differences estimation, covariate balancing, and cross-sample effect calibration, enabling the first systematic correction based on heterogeneity in short-term treatment effects. By bridging randomized experiments and observational datasets, the framework jointly preserves internal validity from experiments and external representativeness from administrative records, overcoming inferential limitations inherent to single-data-source designs. Empirical validation using the STAR randomized experiment and New York State school administrative data demonstrates substantial improvements in both accuracy and external validity of estimated causal effects of class size on academic performance.

Correcting selection bias in observational studies via experimental dataDeveloping a method to weaken assumptions for surrogate estimatorsEstimating treatment effects on primary outcomes using observational and experimental data

From Data Creator to Data Reuser: Distance Matters

Feb 05, 2024
CL
Christine L. Borgman
🏛️ University of California, Los Angeles | University of Amsterdam

This paper addresses the low efficiency of scientific data sharing and reuse by proposing the “Creator–Reuser Distance” theoretical framework—the first systematic identification and modeling of six dimensions impeding knowledge transfer: domain, methodology, collaboration, cataloging, purpose, and time. Departing from conventional technology-centric data delivery paradigms, it reconceptualizes data reuse as a socio-cognitive knowledge exchange process. Drawing on interdisciplinary foundations in scientometrics, information science, and socio-technical systems, the study employs empirically grounded conceptual modeling—not algorithmic or engineering implementation—to uncover how distance constrains knowledge transmission. The findings provide a foundational theory for open science infrastructure development and deliver tiered, actionable intervention strategies tailored to four key stakeholder groups: data creators, reusers, archivists, and funding agencies—thereby enhancing investment efficiency across the data lifecycle and facilitating cross-domain knowledge flow.

Addressing knowledge exchange in data sharingIdentifying dimensions influencing data reuseProviding recommendations for effective data circulation

Latest Papers

What's happening recently
View more

This study addresses the challenge of ensuring rigor in causal inference under multi-source heterogeneous data fusion by proposing a structured design paradigm grounded in the target trial framework. The approach explicitly incorporates the target population and its sampling model into the causal analysis, systematically integrating external controls, generalizability, and transportability assessments through data element alignment, transparent assumption articulation, and emulation of the target trial. Its key innovation lies in anchoring the entire framework to a precise definition of the target population, thereby identifying and mitigating irreconcilable conflicts across data sources. This strategy enhances both the reliability and interpretability of causal conclusions derived from complex, real-world data ecosystems.

causal inferencedata integrationexternal comparator analyses

This study addresses the limitations of individualized clinical decision-making, which is often constrained by the high internal validity but limited external applicability of randomized controlled trials (RCTs) and the strong representativeness yet susceptibility to confounding bias in real-world data (RWD). To overcome these challenges, the authors propose a multi-source data integration paradigm grounded in an explicit causal inference framework. This approach systematically combines RCT and RWD by rigorously defining estimands, ensuring comparability across data sources, and conducting sensitivity analyses. The resulting methodology enhances the reliability and evidentiary strength of treatment effect estimates while providing a practical, regulatory-compliant pathway for generating individualized treatment recommendations.

causal frameworksevidence integrationrandomized controlled trials

Causal Secondary Analysis of Linked Data in the Presence of Mismatch Error

Dec 16, 2025
MS
Martin Slawski
🏛️ University of Virginia

This paper addresses estimation bias in the average treatment effect (ATE) arising from record linkage mismatches in secondary analysis of multi-source observational data. We propose a novel causal inference method that, for the first time, formalizes linkage mismatch as a missing-data mechanism. Our approach constructs an estimating equation framework based on a two-component mixture model and an enhanced EM algorithm, enabling asymptotically consistent ATE estimation and valid statistical inference—even when linkage quality information is unavailable. Simulation studies and empirical applications demonstrate that the method substantially reduces estimation bias, improves ATE consistency, and achieves nominal coverage rates for confidence intervals—outperforming naive analyses that ignore linkage errors. By rigorously accounting for linkage uncertainty, our framework provides a generalizable, statistically principled solution for causal secondary analysis under imperfect record linkage.

Addresses bias from mismatched records in secondary data analysis.Estimates causal effects despite linkage errors in data integration.Proposes a method to handle unknown match status in datasets.

This study addresses long-standing limitations in India’s inter-state migration census data, which have been plagued by uneven state-level coverage and inconsistent measurement practices, leading to systematic biases that undermine analytical reliability. For the first time, the paper systematically disentangles measurement bias from representativeness bias and introduces a data-driven Harmonized Inter-State Migration (HICM) framework. Integrating statistical diagnostics, imputation, smoothing, and bias correction techniques, HICM standardizes and reconciles migration data across states and time. The proposed approach delivers a reproducible, bias-aware preprocessing and validation pipeline that substantially enhances structural consistency and temporal stability. Empirical results demonstrate that the corrected data significantly improve the credibility of migration network analyses, offering policymakers more accurate evidence for informed decision-making.

census datadata harmonizationmeasurement bias

This study addresses the challenge of sample overlap in evidence synthesis from observational studies, which can introduce substantial bias—particularly when individual-level identifiers are unavailable to detect or correct such overlap. To overcome this limitation, the authors propose a novel method grounded in set theory that requires no individual participant data. By encoding the ranges of multiple carefully selected sample characteristics, the approach constructs an index quantifying the degree of sample overlap, enabling inference of overlapping samples and identification of the largest non-overlapping subset. This method fills a critical gap in the secondary use of real-world evidence, where handling sample overlap has been underexplored. Its validity and flexibility are demonstrated across several empirical case studies, significantly enhancing the credibility of synthesized evidence.

biasdata reuseevidence synthesis

Hot Scholars

ML

Maryline Laurent

Telecom SudParis
cybersecurityprivacy enhancing technologiesdigital identityblockchain
BE

Beatriz Esteves

Postdoctoral Researcher, Ghent University
semantic webdata protectionpoliciestrust
DS

Dallas Seitz

Professor, Centre for Addiction and Mental Health
Geriatric Psychiatry
JM

Jared Murray

Associate Professor of Statistics and Machine Learning, University of Texas at Austin
QY

Qishuo Yin

Princeton University
statisticscausal inference