cohort construction

Designing and harmonizing inclusion/exclusion criteria and data selection procedures to assemble reproducible study populations or user cohorts, ensuring they meet validation, serving, and analysis constraints across datasets and segments.

cohortconstruction

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

To address the challenge of achieving FAIR interoperability and reusability for multi-source, heterogeneous scientific data—such as the RADx COVID-19 response data—in large-scale environments, this paper proposes a general-purpose, reproducible, and extensible data harmonization framework. Our approach introduces a novel harmonization paradigm based on parameterized primitive operations and automated execution tracing. It integrates a customizable data representation model, a configurable operation library, and mechanisms for transformation logging and dependency tracking—ensuring protocol reproducibility, process auditability, and transformation reusability. Evaluated in real-world deployment within the RADx Data Hub, the framework significantly lowers the barrier to entry for domain experts, improves harmonization efficiency and transparency, and enables high-quality cross-study analyses. This work provides a scalable, transferable technical pathway for FAIR-compliant data integration across diverse scientific domains.

Addresses challenges in aligning heterogeneous data setsFacilitates data harmonization for FAIR principles complianceSupports reproducible and scalable data integration processes

Grand Challenge: Mediating Between Confirmatory and Exploratory Research Cultures in Health Sciences and Visual Analytics

Aug 15, 2025
VV
Viktor von Wyl
🏛️ University of Zurich | UZH Digital Society Initiative (DSI) | University of Zurich | Department of Informatics

Health sciences—hypothesis-driven and emphasizing reproducibility—clash with visual analytics—iterative, exploratory, and interaction-dependent—leading to cross-disciplinary challenges: terminological misalignment, divergent expectations for data preparation, conflicting validation criteria, and contradictory interpretability requirements. To address this, we propose an integrative framework structured along three dimensions: cultural adaptation, standard harmonization, and process coordination. It specifies seven concrete, actionable steps—the first systematic effort to bridge confirmatory and exploratory research paradigms. Grounded in interdisciplinary co-design, the framework incorporates integrated workflow modeling, a terminology alignment tool, and a multi-stage quality validation benchmark. It enables clinically relevant, reliable, and reproducible collaborative analysis. By fostering deep methodological integration, the framework advances a unified research agenda that enhances scientific rigor, practical feasibility, and clinical translatability of hybrid analytical approaches.

Addressing interdisciplinary challenges like terminology misalignment and validation normsBridging gap between hypothesis-driven health science and exploratory visual analyticsDeveloping frameworks for integrating confirmatory and exploratory research methods

What is Reproducibility in Artificial Intelligence and Machine Learning Research?

Apr 29, 2024
AD
Abhyuday Desai
🏛️ Ready Tensor, Inc. | Georgetown University

The AI/ML community faces a severe reproducibility crisis, primarily driven by conceptual ambiguity in verification terminology—such as “reproducibility,” “replicability,” and “dependency/independence”—which undermines research credibility and scientific progress. To address this, we propose the first five-dimensional verification taxonomy, systematically defining core concepts—including reproducibility, dependency vs. independent re-executability, and direct vs. conceptual replicability—by clarifying their objectives, prerequisites, and evaluation criteria. Our framework integrates conceptual analysis, terminological standardization, and methodological modeling to yield a structured verification guideline. It enhances experimental rigor in study design, fosters consensus across the research community on verification practices, and significantly improves cross-team result reproducibility and outcome reliability.

Addressing the reproducibility crisis in AI/ML researchClarifying validation terminology in AI/ML reproducibilityProviding a framework for validation study design

This study addresses the challenge of inaccurate sample size estimation in nonparametric tolerance limit methods under selection biases—such as weighting bias and censoring—which can induce substantial bias in overall coverage probability. The work extends Scheffé and Tukey’s classical tolerance limit theory to multiple biased sampling scenarios, proposing a novel nonparametric approach based on order statistics. By integrating bias-correction derivations with Monte Carlo simulations, it establishes a robust framework for sample size calculation that accommodates non-representative samples. The method demonstrates strong performance in simulation studies and has been successfully applied to time-to-failure data on Alzheimer’s disease from the Canadian Longitudinal Study on Aging, markedly improving the accuracy and practicality of sample size estimation under biased sampling conditions.

biased samplingprevalent cohortsample size determination

Scientists frequently record experimental metadata in spreadsheets, yet ensuring consistency and standards compliance remains challenging. This paper introduces a spreadsheet-native metadata governance paradigm: customized Excel/CSV templates embed HuBMAP standards; OWL/SKOS ontology-driven controlled vocabularies are integrated; and a web-based real-time semantic validation tool enables immediate, on-entry verification. The approach seamlessly incorporates semantic constraints into familiar spreadsheet workflows—requiring no platform switching or new system adoption. Deployed across the HuBMAP Consortium, it significantly improved multi-omics metadata compliance rates, increased data entry efficiency, and reduced error identification and correction time by over 70%. To our knowledge, this is the first work to deeply embed ontology-based constraints and real-time semantic validation directly within spreadsheet environments, establishing a scalable, practical paradigm for biomedical metadata standardization.

Addressing spreadsheet limitations for consistent experiment-related metadata annotationEnsuring metadata standards compliance in spreadsheet-based scientific data entryProviding quality control for biomedical metadata collection using spreadsheets

Latest Papers

What's happening recently
View more

Data-driven controlled subgroup selection in clinical trials

Dec 17, 2025
MM
Manuel M. Müller
🏛️ University of Cambridge | Novartis Pharma AG | University of Edinburgh | University of Southampton | Nanjing University | Lancaster University

Data-driven subgroup identification in clinical trials suffers from post-selection inference issues, leading to inflated Type I error rates and biased effect estimates—hindering the implementation of precision medicine. To address the dual objective of identifying both *safe subgroups* (with low adverse event risk) and *efficacious subgroups* (with high treatment effect), this paper proposes two novel controlled subgroup selection methods: one based on generalized linear models and another within an isotonic regression framework. For the first time in a regression setting, both methods enable rigorous post-selection inference with guaranteed Type I error control under the null. Comprehensive simulation studies demonstrate robust error rate control across diverse scenarios and quantify sensitivity to modeling assumptions. The proposed methods provide a statistically rigorous, reproducible, and practically applicable toolkit for clinical subgroup analysis.

Addresses post-selection inference to control Type I error ratesDevelops methods for selecting patient subgroups in clinical trialsIdentifies subgroups with high treatment effect or safety from adverse events

This work addresses the inefficiencies in oncology clinical trial statistical workflows—often fragmented, leading to redundant efforts, poor collaboration, and inconsistent analyses—by developing grstat, an open-source R package that integrates standardized analytical tools within a governance framework featuring requirement traceability, peer review, automated testing, and phased validation. By unifying technical implementation with a structured, reproducible process, grstat establishes a shared, auditable, and maintainable analytical toolkit. Empirical application demonstrates that this approach substantially enhances analytical efficiency, consistency, and long-term maintainability, offering academic biostatistics teams a scalable and transferable collaborative paradigm.

clinical trialsoncologyreproducibility

This study addresses the issue of model unfairness toward specific subpopulations in high-stakes clinical settings, stemming from biases in training data. Focusing on intensive care unit (ICU) environments, the authors systematically evaluate how incorporating external electronic health records—such as those from eICU and MIMIC-IV—affects fairness across subgroups. Through multi-source data integration, subgroup performance analysis, post-hoc calibration, and comparative data selection strategies, they demonstrate that simply increasing training data volume does not necessarily improve fairness and may even degrade subgroup performance, thereby challenging the prevailing “more data is better” paradigm. The work proposes a novel framework of joint data-model intervention, showing that combining targeted data augmentation with post-processing calibration effectively enhances both fairness and overall predictive performance.

algorithmic biasdata interventionsdistribution shift

Auditing Reproducibility in Non-Targeted Analysis: 103 LC/GC--HRMS Tools Reveal Temporal Divergence Between Openness and Operability

Dec 23, 2025
SA
Sarah Alsubaie
🏛️ King Abdullah University of Science and Technology (KAUST)

Despite growing adoption of non-targeted analysis (NTA) in food safety, systematic evaluation of liquid/gas chromatography–high-resolution mass spectrometry (LC/GC-HRMS) NTA tools against FAIR principles and the BP4NTA operational pillars—laboratory validation, data/code availability, standardized formats, knowledge integration, and portable implementation—remains lacking. Method: We conducted a longitudinal, systematic audit of 103 NTA tools published between 2004 and 2025, assessing compliance across all six BP4NTA pillars. Contribution/Results: We found a marked increase in openness (56% → 86%) but a paradoxical decline in practical reproducibility (55% → 43%), quantifying for the first time the persistent “discoverable but not runnable” gap. Critical synergistic deficits between Pillar C1 (validation) and C6 (portable implementation) emerged as the primary bottleneck for regulatory-grade reproducibility. This study fills a key gap in food safety NTA tool assessment and proposes a multidimensional audit framework; only 17% of tools satisfy both validation and portability criteria—providing empirical grounding and actionable pathways toward fully reproducible NTA workflows.

Assessing divergence between data openness and operational portabilityAuditing reproducibility of non-targeted analysis toolsIdentifying gaps in validation, standardization, and implementation

Reproducibility remains a central challenge in computational social science, where complex workflows, evolving software ecosystems, and inconsistent documentation hinder researchers ability to re-execute published methods. This study presents a systematic evaluation of reproducibility across three conditions: uncurated documentation, curated documentation, and curated documentation paired with a preset execution environment. Using 47 usability test sessions, we combine behavioral performance indicators (success rates, task time, and error profiles) with questionnaire data and thematic analysis to identify technical and conceptual barriers to reproducibility. Curated documentation substantially reduced repository-level errors and improved users ability to interpret method outputs. Standardizing the execution environment further improved reproducibility, yielding the highest success rate and shortest task completion times. Across conditions, participants frequently relied on AI tools for troubleshooting, often enabling independent resolution of issues without facilitator intervention. Our findings demonstrate that reproducibility barriers are multi-layered and require coordinated improvements in documentation quality, environment stability, and conceptual clarity. We discuss implications for the design of reproducibility platforms and infrastructure in computational social science.

computational social sciencedocumentationexecution environment

Hot Scholars

XS

Xu Shi

University of Michigan
Electronic Health RecordCausal InferenceNegative ControlMachine Translation
LW

Lingfei Wu

University of Pittsburgh
science of scienceteam science
DZ

Doudou Zhou

National University of Singapore
High-dimensional StatisticsEHR Data AnalysisChange-point DetectionTransfer Learning
TA

Tawfiq Ammari

Assistant Professor, Rutgers University School of Communication and Information
Data ScienceHuman-Computer InteractionCSCWSTS
TC

Tianxi Cai

Harvard University
statisticsbiostatisticsmodelingprediction