Score
Methods for linking, validating, harmonizing, and merging administrative records and survey data across sources and time to construct analysis-ready measures for causal and descriptive studies.
This study addresses the challenge of performing effective statistical inference on vertically partitioned health data across multiple institutions without sharing individual-level private information. Through a scoping review incorporating interdisciplinary database searches, systematic feature extraction, and citation tracking, 30 relevant studies were identified and evaluated. The findings reveal that existing approaches predominantly focus on linear and logistic regression models, yet commonly lack rigorous validation of equivalence to centralized analyses, incur high communication overhead, and rarely offer quantifiable privacy guarantees. This work provides the first systematic synthesis of methods in this domain, highlighting critical gaps concerning analytical equivalence, communication efficiency, and formal privacy protection, thereby establishing a foundational framework and guiding future research directions.
A significant gap exists between theoretical definitions of data quality dimensions—such as accuracy, completeness, consistency, and timeliness—as stipulated in standards (e.g., ISO/IEC 25012) and the actual functionalities implemented in widely used data quality tools, with no systematic mapping analysis to date. Method: We conducted a systematic literature review and performed functional reverse engineering on seven mainstream open-source data quality tools. Contribution/Results: We present the first many-to-many mapping framework linking data quality dimensions to concrete tool capabilities, introducing a cross-dimensional, fine-grained functional categorization schema and generating a structured correspondence matrix covering all core dimensions. This work bridges the theory–practice divide, delivers an actionable guide for data quality assessment, and substantially enhances the scientific rigor and reusability of tool selection, functionality design, and standard implementation.
This study addresses the entity linking problem in multi-source heterogeneous data lacking unique identifiers. We propose a probabilistic record linkage method that balances accuracy and scalability. Methodologically, we introduce the Stochastic EM algorithm into latent-variable generative models for the first time, explicitly modeling dependencies among link decisions and enforcing one-to-one constraints, while enabling robust linking under variable-quality fields. Our approach innovatively supports dynamic precision–efficiency trade-offs, effectively handling real-world challenges such as information evolution, data entry errors, and low-quality attributes. Extensive evaluation on large-scale real-world healthcare data demonstrates high linkage accuracy; simulation experiments confirm strong robustness to noise and missing values. The open-source R package FlexRL has been released and deployed in production environments.
This study addresses the inherent differences in measurement mechanisms and sample representativeness between intelligent surveys (e.g., sensor- and AI-driven data collection) and traditional surveys. Methodologically, it distinguishes two integration paradigms: “hybrid-mode” integration—emphasizing outcome alignment to enable direct data merging—and “multi-source integration”—fusing heterogeneous data during modeling to leverage complementary strengths. A context-driven decision framework is proposed, integrating sensing technologies, machine learning, and statistical modeling for synergistic analysis. The key contribution lies in the first systematic delineation of the applicability boundaries and trade-offs between these two paradigms, offering actionable strategies for official statistics integration. Empirical application demonstrates substantial improvements in data quality and timeliness for dynamic behavioral surveys—particularly travel behavior—thereby advancing theoretical foundations and practical guidelines for statistical modernization in multi-source data environments. (149 words)
To address the challenge of achieving FAIR interoperability and reusability for multi-source, heterogeneous scientific data—such as the RADx COVID-19 response data—in large-scale environments, this paper proposes a general-purpose, reproducible, and extensible data harmonization framework. Our approach introduces a novel harmonization paradigm based on parameterized primitive operations and automated execution tracing. It integrates a customizable data representation model, a configurable operation library, and mechanisms for transformation logging and dependency tracking—ensuring protocol reproducibility, process auditability, and transformation reusability. Evaluated in real-world deployment within the RADx Data Hub, the framework significantly lowers the barrier to entry for domain experts, improves harmonization efficiency and transparency, and enables high-quality cross-study analyses. This work provides a scalable, transferable technical pathway for FAIR-compliant data integration across diverse scientific domains.
This study addresses selection bias in estimating long-term causal effects—such as graduation rates—from observational studies. We propose a novel control function approach that leverages experimental estimates of treatment effects on short-term outcomes (e.g., eighth-grade test scores) to correct for unobserved confounding in large-scale administrative observational data. Our method integrates insights from difference-in-differences estimation, covariate balancing, and cross-sample effect calibration, enabling the first systematic correction based on heterogeneity in short-term treatment effects. By bridging randomized experiments and observational datasets, the framework jointly preserves internal validity from experiments and external representativeness from administrative records, overcoming inferential limitations inherent to single-data-source designs. Empirical validation using the STAR randomized experiment and New York State school administrative data demonstrates substantial improvements in both accuracy and external validity of estimated causal effects of class size on academic performance.
This paper addresses the low efficiency of scientific data sharing and reuse by proposing the “Creator–Reuser Distance” theoretical framework—the first systematic identification and modeling of six dimensions impeding knowledge transfer: domain, methodology, collaboration, cataloging, purpose, and time. Departing from conventional technology-centric data delivery paradigms, it reconceptualizes data reuse as a socio-cognitive knowledge exchange process. Drawing on interdisciplinary foundations in scientometrics, information science, and socio-technical systems, the study employs empirically grounded conceptual modeling—not algorithmic or engineering implementation—to uncover how distance constrains knowledge transmission. The findings provide a foundational theory for open science infrastructure development and deliver tiered, actionable intervention strategies tailored to four key stakeholder groups: data creators, reusers, archivists, and funding agencies—thereby enhancing investment efficiency across the data lifecycle and facilitating cross-domain knowledge flow.
This study addresses the challenge of ensuring rigor in causal inference under multi-source heterogeneous data fusion by proposing a structured design paradigm grounded in the target trial framework. The approach explicitly incorporates the target population and its sampling model into the causal analysis, systematically integrating external controls, generalizability, and transportability assessments through data element alignment, transparent assumption articulation, and emulation of the target trial. Its key innovation lies in anchoring the entire framework to a precise definition of the target population, thereby identifying and mitigating irreconcilable conflicts across data sources. This strategy enhances both the reliability and interpretability of causal conclusions derived from complex, real-world data ecosystems.
This study addresses the limitations of individualized clinical decision-making, which is often constrained by the high internal validity but limited external applicability of randomized controlled trials (RCTs) and the strong representativeness yet susceptibility to confounding bias in real-world data (RWD). To overcome these challenges, the authors propose a multi-source data integration paradigm grounded in an explicit causal inference framework. This approach systematically combines RCT and RWD by rigorously defining estimands, ensuring comparability across data sources, and conducting sensitivity analyses. The resulting methodology enhances the reliability and evidentiary strength of treatment effect estimates while providing a practical, regulatory-compliant pathway for generating individualized treatment recommendations.
This paper addresses estimation bias in the average treatment effect (ATE) arising from record linkage mismatches in secondary analysis of multi-source observational data. We propose a novel causal inference method that, for the first time, formalizes linkage mismatch as a missing-data mechanism. Our approach constructs an estimating equation framework based on a two-component mixture model and an enhanced EM algorithm, enabling asymptotically consistent ATE estimation and valid statistical inference—even when linkage quality information is unavailable. Simulation studies and empirical applications demonstrate that the method substantially reduces estimation bias, improves ATE consistency, and achieves nominal coverage rates for confidence intervals—outperforming naive analyses that ignore linkage errors. By rigorously accounting for linkage uncertainty, our framework provides a generalizable, statistically principled solution for causal secondary analysis under imperfect record linkage.
This study addresses long-standing limitations in India’s inter-state migration census data, which have been plagued by uneven state-level coverage and inconsistent measurement practices, leading to systematic biases that undermine analytical reliability. For the first time, the paper systematically disentangles measurement bias from representativeness bias and introduces a data-driven Harmonized Inter-State Migration (HICM) framework. Integrating statistical diagnostics, imputation, smoothing, and bias correction techniques, HICM standardizes and reconciles migration data across states and time. The proposed approach delivers a reproducible, bias-aware preprocessing and validation pipeline that substantially enhances structural consistency and temporal stability. Empirical results demonstrate that the corrected data significantly improve the credibility of migration network analyses, offering policymakers more accurate evidence for informed decision-making.
This study addresses the challenge of sample overlap in evidence synthesis from observational studies, which can introduce substantial bias—particularly when individual-level identifiers are unavailable to detect or correct such overlap. To overcome this limitation, the authors propose a novel method grounded in set theory that requires no individual participant data. By encoding the ranges of multiple carefully selected sample characteristics, the approach constructs an index quantifying the degree of sample overlap, enabling inference of overlapping samples and identification of the largest non-overlapping subset. This method fills a critical gap in the secondary use of real-world evidence, where handling sample overlap has been underexplored. Its validity and flexibility are demonstrated across several empirical case studies, significantly enhancing the credibility of synthesized evidence.