epidemiological data preprocessing

Designs and implements reproducible workflows to clean, harmonize, and prepare epidemiological and epidemic surveillance data for analysis or modeling. This includes error and duplicate detection, missing-data imputation, alignment and aggregation of time series, adjustment for reporting delays and changing case definitions, standardization of formats and geocodes, feature extraction and metadata/provenance capture, and privacy-preserving de‑identification to produce documented datasets ready for statistical or mechanistic use.

epidemiologicaldatapreprocessing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address the challenge of achieving FAIR interoperability and reusability for multi-source, heterogeneous scientific data—such as the RADx COVID-19 response data—in large-scale environments, this paper proposes a general-purpose, reproducible, and extensible data harmonization framework. Our approach introduces a novel harmonization paradigm based on parameterized primitive operations and automated execution tracing. It integrates a customizable data representation model, a configurable operation library, and mechanisms for transformation logging and dependency tracking—ensuring protocol reproducibility, process auditability, and transformation reusability. Evaluated in real-world deployment within the RADx Data Hub, the framework significantly lowers the barrier to entry for domain experts, improves harmonization efficiency and transparency, and enables high-quality cross-study analyses. This work provides a scalable, transferable technical pathway for FAIR-compliant data integration across diverse scientific domains.

Addresses challenges in aligning heterogeneous data setsFacilitates data harmonization for FAIR principles complianceSupports reproducible and scalable data integration processes

Harmonizing Community Science Datasets to Model Highly Pathogenic Avian Influenza (HPAI) in Birds in the Subantarctic

Dec 07, 2025
RL
Richard Littauer
🏛️ Te Herenga Waka Victoria University of Wellington

Assessing the impact of highly pathogenic avian influenza (HPAI) on bird populations in the sub-Antarctic faces critical challenges due to high heterogeneity, poor standardization, and variable quality of community science data. Method: We developed a reproducible workflow integrating multi-source observational data from eBird, iNaturalist, and GBIF. We introduced a novel data fusion and cleaning framework for wildlife disease modeling, enabling spatiotemporal alignment, species identification verification, and outlier control across platforms. Contribution/Results: This study produced the first comprehensive, sub-Antarctic–wide avian mortality dataset covering major islands. Leveraging ecological niche modeling and epidemiological extrapolation, we predicted HPAI-associated fatality risk for unobserved species. Our analysis yielded revised population trajectory and mortality estimates for 12 key avian species under HPAI outbreaks, substantially enhancing capacity for dynamic monitoring and early warning of wildlife disease outbreaks in remote regions.

Harmonizing heterogeneous community science datasets for epidemiologyModeling HPAI impact on subantarctic bird populations using aggregated dataPredicting unknown bird demographics and mortality rates from HPAI

Monitoring a developing pandemic with available data

Aug 19, 2023
ML
María Luz Gámiz
🏛️ University of Granada | Heidelberg University | Bayes Business School, City, University of London | University of Graz

In developing epidemic surveillance systems, daily reported data on infections, hospitalizations, deaths, and recoveries often suffer from missingness, inconsistency, and reporting delays. To address this, we propose a dynamic Bayesian statistical modeling framework. Methodologically: (1) we explicitly model the missing-data mechanism, disentangling reporting bias from the underlying epidemiological process; (2) we integrate calendar effects and domain-informed epidemiological priors to enable structured incorporation of expert knowledge; and (3) we build an updateable system for real-time estimation and short-term forecasting. Empirical evaluation on multi-source French COVID-19 data demonstrates substantial improvements in real-time estimation accuracy of key indicators—including the effective reproduction number, hospital burden, and mortality risk—as well as 7–14-day forecast reliability. This work establishes the first reusable, interpretable, and robust statistical modeling benchmark for public health response.

Addressing novel missing data challenges with statistically rigorous solutionsIncorporating calendar effects and expert knowledge into dynamic pandemic forecastingModeling pandemic severity indicators using incomplete daily reported data

AERO: An autonomous platform for continuous research

May 23, 2025
VH
Valérie Hayot-Sasson
🏛️ University of Chicago | Argonne National Laboratory | Globus | University Carlos III of Madrid

The COVID-19 pandemic revealed critical deficiencies in conventional public health data platforms—particularly regarding automation, continuity, and cross-sector collaboration. To address these gaps, we propose and implement an autonomous platform designed for continuous research, featuring the first end-to-end automated closed-loop architecture that integrates dynamic data governance and multi-stakeholder collaborative governance. The platform leverages Globus for secure, trusted data transfer and identity management, and GitHub for workflow versioning and CI/CD-driven automated execution. It supports fully automated acquisition, validation, transformation, analysis, and policy-governed sharing of surveillance data. Deployed in two real-world public health monitoring scenarios, the system demonstrates operational efficacy; scalability is validated via synthetic workload benchmarking. All system designs, source code, and experimental resources are openly released to ensure full reproducibility.

Automates data ingestion, validation, and analysis for epidemiologyDevelops AERO platform for continuous pandemic data collaborationEnables secure multi-entity data sharing using Globus and GitHub

Low quality exposure and point processes with a view to the first phase of a pandemic (preprint)

Aug 19, 2023
ML
María Luz Gámiz
🏛️ University of Granada | Heidelberg University | City, University of London

During the early phase of epidemics, modeling is hindered by ambiguously defined exposure variables (e.g., infection counts), severe data scarcity, and poor data quality. Method: This paper proposes a point-process modeling framework for “low-quality exposure,” based on an inhomogeneous Poisson process that incorporates a time-varying exposure function and robust statistical estimation—enabling dynamic evolution of exposure definitions over time without reliance on high-fidelity epidemiological parameters. Contribution/Results: We formally define “low-quality exposure” for the first time and develop a lightweight, real-time deployable cross-national early-warning model. Evaluated on early-phase French COVID-19 hospitalization and infection data, the method significantly improves short-term forecasting stability and interpretability of daily new hospitalizations, while demonstrating strong cross-country adaptability.

Analyzing pandemic data with poorly defined exposure measurementsDeveloping methodology for changing exposure definitions over timeForecasting hospitalizations using low-quality infection data

Latest Papers

What's happening recently
View more

This study addresses longstanding challenges in wastewater surveillance—namely data fragmentation, inconsistent metadata, and insufficient interoperability—by proposing an open, relational, and modular data model and interoperability framework. Designed to align with established standards such as PHA4GE and the U.S. CDC’s National Wastewater Surveillance System, the framework supports transparent and ethical data use under FAIR principles. It introduces novel features including public health action tables, links to external databases (e.g., GISAID, GenBank), and documentation of analytical workflows, while enhancing multidimensional relational modeling and facilitating conversion between wide and long data formats. Deployed across 25 countries, the framework has been adopted by the Public Health Agency of Canada and adapted by the European Union’s Sewage Sentinel System, demonstrating significant advantages over six existing wastewater data standards across 25 key criteria.

data interoperabilityfragmented data systemsmetadata standardization

This study addresses the systemic gaps in public health emergency response revealed by the COVID-19 pandemic, particularly the absence of a comprehensive, data-driven risk management framework. To bridge this gap, the work proposes the first adaptation of the industrial DMAIC (Define, Measure, Analyze, Improve, Control) systems informatics paradigm to epidemic control. By integrating medical diagnostics, statistical sampling, data visualization, artificial intelligence, telemedicine, and resource optimization, the framework establishes an intelligent decision-support system spanning the entire epidemic lifecycle. This approach not only enhances the resilience and responsiveness of health systems but also fosters interdisciplinary convergence between systems informatics and public health, offering a replicable and scalable methodological foundation for managing future public health emergencies.

DMAICEpidemic InformaticsEpidemic Response

This work proposes an AI agent–driven workflow to address the high costs of reproducing large-scale empirical studies, which often stem from discrepancies in computational environments, code, and documentation. The approach decouples scientific reasoning from computational execution: researchers supply standardized diagnostic templates, and the system automatically retrieves and orchestrates reproduction materials within a version-controlled environment. A structured knowledge layer captures failure patterns, enabling adaptive reproduction across heterogeneous studies while ensuring transparency and stability of the analytical pipeline. Evaluated on 92 instrumental variable studies, the method achieves an 87% end-to-end reproduction success rate; when data and code are available, it attains 100% success at both the paper and model levels.

empirical dataexecution bottlenecklarge-scale reanalysis

This study addresses the high cost and prolonged turnaround time of whole-genome sequencing, which hinder its routine use for hospital outbreak surveillance in resource-limited settings. To overcome this limitation, the authors propose a novel, actionable tiered monitoring framework that systematically integrates three readily available clinical data streams—MALDI-TOF mass spectrometry profiles, antimicrobial resistance phenotypes, and electronic health records—into a machine learning model for rapid and cost-effective outbreak detection. The approach substantially reduces reliance on genomic sequencing while significantly enhancing detection performance across multiple pathogen species. Moreover, the multimodal fusion strategy precisely identifies high-risk transmission routes associated with invasive procedures and frequent clinical workflows, thereby offering targeted opportunities for infection control interventions.

hospital surveillanceinfection controlmultimodal data

This study addresses the limited research credibility of existing synthetic electronic health records, which often suffer from inconsistencies between clinical processes and observations. The authors propose a two-stage integrated pipeline: first, a knowledge-guided generative model simulates high-fidelity patient trajectories by modeling nearly 32,000 clinical events; second, a large language model–based automated auditing module detects clinical contradictions, such as contraindicated medication prescriptions. Evaluated on 18,071 synthetic records, the method achieves high statistical fidelity (R² = 0.99), substantially reduces clinical inconsistencies, and enables downstream task performance comparable to or better than that achieved with real data—all without privacy leakage risks (F1 = 0.51). This work represents the first integration of knowledge-guided generation with LLM-driven clinical consistency auditing, significantly enhancing the clinical plausibility of synthetic medical records.

clinical consistencydata fidelityhealthcare data privacy

Hot Scholars

TC

Tanujit Chakraborty

Associate Professor of Statistics and Data Science at Sorbonne University
Machine LearningTime Series ForecastingSpatial StatisticsHealth Data Science
AW

Ander Wilson

Associate Professor of Statistics, Colorado State University
BiostatisticsData ScienceEnvironmental StatisticsEnvironmental Health
JP

Junyong Park

Seoul National University
Computer Architecture
JA

Julia A. Palacios

Stanford University
PhylodynamicsEvolutionary GeneticsEpidemiology and Infectious DiseasesBayesian Nonparametrics