Score
Designs and implements reproducible workflows to clean, harmonize, and prepare epidemiological and epidemic surveillance data for analysis or modeling. This includes error and duplicate detection, missing-data imputation, alignment and aggregation of time series, adjustment for reporting delays and changing case definitions, standardization of formats and geocodes, feature extraction and metadata/provenance capture, and privacy-preserving de‑identification to produce documented datasets ready for statistical or mechanistic use.
To address the challenge of achieving FAIR interoperability and reusability for multi-source, heterogeneous scientific data—such as the RADx COVID-19 response data—in large-scale environments, this paper proposes a general-purpose, reproducible, and extensible data harmonization framework. Our approach introduces a novel harmonization paradigm based on parameterized primitive operations and automated execution tracing. It integrates a customizable data representation model, a configurable operation library, and mechanisms for transformation logging and dependency tracking—ensuring protocol reproducibility, process auditability, and transformation reusability. Evaluated in real-world deployment within the RADx Data Hub, the framework significantly lowers the barrier to entry for domain experts, improves harmonization efficiency and transparency, and enables high-quality cross-study analyses. This work provides a scalable, transferable technical pathway for FAIR-compliant data integration across diverse scientific domains.
Assessing the impact of highly pathogenic avian influenza (HPAI) on bird populations in the sub-Antarctic faces critical challenges due to high heterogeneity, poor standardization, and variable quality of community science data. Method: We developed a reproducible workflow integrating multi-source observational data from eBird, iNaturalist, and GBIF. We introduced a novel data fusion and cleaning framework for wildlife disease modeling, enabling spatiotemporal alignment, species identification verification, and outlier control across platforms. Contribution/Results: This study produced the first comprehensive, sub-Antarctic–wide avian mortality dataset covering major islands. Leveraging ecological niche modeling and epidemiological extrapolation, we predicted HPAI-associated fatality risk for unobserved species. Our analysis yielded revised population trajectory and mortality estimates for 12 key avian species under HPAI outbreaks, substantially enhancing capacity for dynamic monitoring and early warning of wildlife disease outbreaks in remote regions.
In developing epidemic surveillance systems, daily reported data on infections, hospitalizations, deaths, and recoveries often suffer from missingness, inconsistency, and reporting delays. To address this, we propose a dynamic Bayesian statistical modeling framework. Methodologically: (1) we explicitly model the missing-data mechanism, disentangling reporting bias from the underlying epidemiological process; (2) we integrate calendar effects and domain-informed epidemiological priors to enable structured incorporation of expert knowledge; and (3) we build an updateable system for real-time estimation and short-term forecasting. Empirical evaluation on multi-source French COVID-19 data demonstrates substantial improvements in real-time estimation accuracy of key indicators—including the effective reproduction number, hospital burden, and mortality risk—as well as 7–14-day forecast reliability. This work establishes the first reusable, interpretable, and robust statistical modeling benchmark for public health response.
The COVID-19 pandemic revealed critical deficiencies in conventional public health data platforms—particularly regarding automation, continuity, and cross-sector collaboration. To address these gaps, we propose and implement an autonomous platform designed for continuous research, featuring the first end-to-end automated closed-loop architecture that integrates dynamic data governance and multi-stakeholder collaborative governance. The platform leverages Globus for secure, trusted data transfer and identity management, and GitHub for workflow versioning and CI/CD-driven automated execution. It supports fully automated acquisition, validation, transformation, analysis, and policy-governed sharing of surveillance data. Deployed in two real-world public health monitoring scenarios, the system demonstrates operational efficacy; scalability is validated via synthetic workload benchmarking. All system designs, source code, and experimental resources are openly released to ensure full reproducibility.
During the early phase of epidemics, modeling is hindered by ambiguously defined exposure variables (e.g., infection counts), severe data scarcity, and poor data quality. Method: This paper proposes a point-process modeling framework for “low-quality exposure,” based on an inhomogeneous Poisson process that incorporates a time-varying exposure function and robust statistical estimation—enabling dynamic evolution of exposure definitions over time without reliance on high-fidelity epidemiological parameters. Contribution/Results: We formally define “low-quality exposure” for the first time and develop a lightweight, real-time deployable cross-national early-warning model. Evaluated on early-phase French COVID-19 hospitalization and infection data, the method significantly improves short-term forecasting stability and interpretability of daily new hospitalizations, while demonstrating strong cross-country adaptability.
This study addresses longstanding challenges in wastewater surveillance—namely data fragmentation, inconsistent metadata, and insufficient interoperability—by proposing an open, relational, and modular data model and interoperability framework. Designed to align with established standards such as PHA4GE and the U.S. CDC’s National Wastewater Surveillance System, the framework supports transparent and ethical data use under FAIR principles. It introduces novel features including public health action tables, links to external databases (e.g., GISAID, GenBank), and documentation of analytical workflows, while enhancing multidimensional relational modeling and facilitating conversion between wide and long data formats. Deployed across 25 countries, the framework has been adopted by the Public Health Agency of Canada and adapted by the European Union’s Sewage Sentinel System, demonstrating significant advantages over six existing wastewater data standards across 25 key criteria.
This study addresses the systemic gaps in public health emergency response revealed by the COVID-19 pandemic, particularly the absence of a comprehensive, data-driven risk management framework. To bridge this gap, the work proposes the first adaptation of the industrial DMAIC (Define, Measure, Analyze, Improve, Control) systems informatics paradigm to epidemic control. By integrating medical diagnostics, statistical sampling, data visualization, artificial intelligence, telemedicine, and resource optimization, the framework establishes an intelligent decision-support system spanning the entire epidemic lifecycle. This approach not only enhances the resilience and responsiveness of health systems but also fosters interdisciplinary convergence between systems informatics and public health, offering a replicable and scalable methodological foundation for managing future public health emergencies.
This work proposes an AI agent–driven workflow to address the high costs of reproducing large-scale empirical studies, which often stem from discrepancies in computational environments, code, and documentation. The approach decouples scientific reasoning from computational execution: researchers supply standardized diagnostic templates, and the system automatically retrieves and orchestrates reproduction materials within a version-controlled environment. A structured knowledge layer captures failure patterns, enabling adaptive reproduction across heterogeneous studies while ensuring transparency and stability of the analytical pipeline. Evaluated on 92 instrumental variable studies, the method achieves an 87% end-to-end reproduction success rate; when data and code are available, it attains 100% success at both the paper and model levels.
This study addresses the high cost and prolonged turnaround time of whole-genome sequencing, which hinder its routine use for hospital outbreak surveillance in resource-limited settings. To overcome this limitation, the authors propose a novel, actionable tiered monitoring framework that systematically integrates three readily available clinical data streams—MALDI-TOF mass spectrometry profiles, antimicrobial resistance phenotypes, and electronic health records—into a machine learning model for rapid and cost-effective outbreak detection. The approach substantially reduces reliance on genomic sequencing while significantly enhancing detection performance across multiple pathogen species. Moreover, the multimodal fusion strategy precisely identifies high-risk transmission routes associated with invasive procedures and frequent clinical workflows, thereby offering targeted opportunities for infection control interventions.
This study addresses the limited research credibility of existing synthetic electronic health records, which often suffer from inconsistencies between clinical processes and observations. The authors propose a two-stage integrated pipeline: first, a knowledge-guided generative model simulates high-fidelity patient trajectories by modeling nearly 32,000 clinical events; second, a large language model–based automated auditing module detects clinical contradictions, such as contraindicated medication prescriptions. Evaluated on 18,071 synthetic records, the method achieves high statistical fidelity (R² = 0.99), substantially reduces clinical inconsistencies, and enables downstream task performance comparable to or better than that achieved with real data—all without privacy leakage risks (F1 = 0.51). This work represents the first integration of knowledge-guided generation with LLM-driven clinical consistency auditing, significantly enhancing the clinical plausibility of synthetic medical records.