curate clinical datasets

Designs and builds curated clinical datasets and associated query collections by collecting representative medical records and images, annotating them with clinical labels and multi-specialty metadata, cleaning and correcting systematic errors, and documenting data provenance and curation protocols. Also integrates external benchmark sets and prepares blinded evaluation samples and real‑world query datasets (including clinician‑generated queries) to support reproducible, robust evaluation.

curateclinicaldatasets

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.27
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

In the Picture: Medical Imaging Datasets, Artifacts, and their Living Review

Jan 18, 2025
AJ
Amelia Jim'enez-S'anchez
🏛️ IT University of Copenhagen | University of Copenhagen | Radboud University Medical Center | Universitat de Barcelona | Technical University of Denmark | CONICET | University of Buenos Aires | Emory University | Stanford University | University of Groningen | Aarhus University | The Hebrew University of Jerusalem | University of Southern Denmark | Lunit | Cerebriu A/S | Federal University of Espírito Santo | German Cancer Research Center | Heidelberg University | University of Bern | Plain Medical | Oxford

Medical imaging datasets commonly suffer from label noise, shortcut learning, missing metadata, and challenges in retrospectively addressing newly discovered issues (e.g., biases, artifacts) post-publication—undermining model robustness and clinical reliability. To address these challenges, we propose the first “dynamic living review” paradigm for medical imaging datasets, establishing a full-lifecycle data governance system. We design a structured SQL database and a standardized metadata framework to enable traceable, cross-referenced linkage among datasets, publications, and documented research flaws (e.g., biases, annotation errors, shortcut effects). Additionally, we develop an open-source, web-based interactive knowledge graph to facilitate community-driven verification and iterative curation. The system has archived over 100 documented flaws across multimodal imaging datasets, advancing practical adoption of standardized data documentation, annotation quality assessment, and fairness auditing in medical AI.

Algorithm PerformanceDataset QualityMedical Image Analysis

Technical specification of a framework for the collection of clinical images and data

Jul 29, 2025
AM
Alistair Mackenzie
🏛️ Royal Surrey NHS Foundation Trust | Dutch Expert Centre for Screening (LRCB)

This study addresses critical challenges in clinical AI development—namely, the difficulty of acquiring high-quality data, poor timeliness, insufficient multi-center collaboration, and elevated ethical and regulatory risks. We propose a sustainably updated, multi-center clinical imaging and data automation acquisition framework. The framework integrates a real-time streaming acquisition system, a dynamic information governance protocol, a privacy-preserving data-sharing infrastructure, and a full-lifecycle ethics review mechanism, concurrently incorporating both historical and real-time clinical data. Its key innovation lies in unifying data timeliness, representativeness, and regulatory compliance: it enables secure cross-institutional collaboration while ensuring GDPR and HIPAA compliance, thereby significantly enhancing the generalizability and validation reliability of AI models in real-world clinical settings. Experimental results demonstrate that the resulting dataset improves downstream model performance by 12.3% and reduces data update latency to under four hours.

Automated vs. semi-automated methods for dataset freshnessEnsuring ethical, safe data collection and sharing processesFramework for collecting clinical images and AI training data

CURE: A dataset for Clinical Understanding & Retrieval Evaluation

Dec 09, 2024
NS
Nadia Sheikh
🏛️ Clinia | University of Waterloo

Existing dense retrievers exhibit poor out-of-distribution generalization in clinical settings and lack multilingual evaluation benchmarks tailored to medical decision-making. To address this, we introduce CURE—the first point-of-care–oriented, multilingual clinical retrieval dataset. CURE spans 10 medical domains and comprises 2,000 real-world clinical queries, supporting both monolingual (English) and cross-lingual (French/Spanish → English) evaluation. It was co-constructed by clinical experts using an ad-hoc paradigm combining manual annotation with expert verification, and is compatible with both dense and sparse retrieval models under standard IR metrics (MRR, NDCG@10). The dataset and its open-source implementation are integrated into the Hugging Face ecosystem. Empirical evaluation reveals substantial performance degradation of state-of-the-art dense retrievers on CURE, exposing critical limitations in clinical domain generalization. CURE has emerged as a new de facto standard for evaluating medical AI retrieval systems.

Addressing generalization issues in dense retrievers for medical queriesLack of domain-specific test sets for clinical retrieval evaluationNeed for multilingual medical retrieval datasets in healthcare

Datasheets for Healthcare AI: A Framework for Transparency and Bias Mitigation

Jan 09, 2025
MS
Marjia Siddik
🏛️ ADAPT Centre | Dublin City University

Medical AI faces severe challenges related to clinical unfairness arising from biased, incomplete, or non-auditable training data. To address this, we propose the first standardized, machine-readable datasheet framework specifically designed for medical AI. The framework integrates FAIR principles (Findable, Accessible, Interoperable, Reusable) with JSON-LD–structured metadata and embeds regulatory compliance checks alongside an automated bias-risk assessment rule engine. It significantly enhances data transparency and auditability: in pilot deployments across three hospitals, it systematically identified seven categories of latent bias sources and successfully supported two FDA-submitted AI products in passing data governance reviews. This work establishes a scalable, reproducible methodology and practical paradigm for trustworthy data governance in medical AI—bridging technical rigor, regulatory requirements, and ethical accountability.

Data BiasError CorrectionIncomplete Data

Existing clinical reasoning evaluation benchmarks predominantly rely on unstructured or static data, failing to capture the structured and interoperable nature of real-world electronic health records (EHRs). To address this gap, this work proposes a novel pipeline that integrates staged large language model (LLM) generation with terminology-anchored validation and repair, yielding MedCase-Structured—the first HL7 FHIR R4–compliant structured dataset for clinical reasoning assessment. Built upon the MedCaseReasoning benchmark, the pipeline successfully generates valid FHIR bundles for 82.5% of cases. Experimental results demonstrate that LLMs exhibit significantly lower diagnostic accuracy when provided with structured FHIR inputs compared to plain text, underscoring both the necessity of aligning evaluations with authentic clinical workflows and the innovative contribution of this dataset.

benchmarkingclinical reasoningelectronic health records

Latest Papers

What's happening recently
View more

This study addresses the critical bottleneck in developing general-purpose medical foundation models—namely, the scarcity of large-scale, standardized, and high-quality medical imaging datasets. To tackle this challenge, the authors conduct a systematic survey of over 1,000 open-source medical imaging datasets and construct the first comprehensive landscape encompassing multiple imaging modalities, clinical tasks, and anatomical regions. They propose a metadata-driven fusion paradigm (MDFP) to structurally integrate these fragmented resources. The work delivers an interactive data discovery portal, a unified dataset catalog, and a scalable, highly reusable structured repository, substantially enhancing the discoverability and utilization efficiency of medical imaging data and thereby establishing a robust data foundation for future medical foundation model research.

data fragmentationdataset scarcityfoundation models

This study addresses the challenge of unreliable AI models and diminished clinical trust stemming from opaque data quality reporting in the secondary use of electronic health records (EHRs). To this end, the authors propose the first comprehensive framework for transparent data quality reporting across the entire EHR lifecycle. The framework innovatively distinguishes between data producers and consumers, explicitly defines five critical phases, and maps established data quality dimensions to specific workflow stages. Through iterative stakeholder and process analysis, a structured reporting mechanism is developed and validated on real-world datasets, demonstrating its ability to effectively trace the origins of data quality issues. The approach significantly enhances data interpretability, fitness-for-use, and governance efficacy, thereby providing a robust foundation for trustworthy AI development and clinical research.

clinical AIdata lifecycledata quality

This study addresses the challenges of accurately querying structured data and extracting information from unstructured clinical text in electronic health records (EHRs). To this end, the authors propose a unified framework that integrates large language models (LLMs) with retrieval-augmented generation (RAG): LLMs are employed to execute structured queries (e.g., Pandas operations), while RAG enhances information extraction from unstructured clinical narratives. The work introduces an innovative automatic evaluation pipeline based on synthetically generated question-answer pairs, combining exact match metrics, semantic similarity scores, and human assessments. Evaluated on a subset of MIMIC-III, the approach demonstrates improved semantic accuracy and task adaptability, offering clinical data science a flexible and reliable tool for automated reasoning and evaluation.

Clinical Data ScienceElectronic Health RecordsInformation Extraction

This work addresses the lack of a unified, machine-verifiable data specification in medical imaging AI, which hinders consistent dataset structure, annotation provenance, quality documentation, and ML-readiness. To bridge this gap, we propose VIDS—an open standard that, for the first time, integrates standardized folder organization, naming conventions, annotation provenance schemas, and quality documentation within a single framework, along with 21 machine-verifiable rules. Built around the NIfTI working format while preserving DICOM metadata, VIDS provides an open-source validator installable via PyPI and supports export to mainstream frameworks such as nnU-Net, MONAI, and COCO. Evaluation reveals that four widely used datasets comply with only 20–39% of VIDS dimensions. We also release LIDC-Hybrid-100, a fully compliant reference dataset comprising 100 CT scans annotated by consensus among four radiologists (mean Dice: 0.7765), which passes all 21 validation checks.

annotated datasetsannotation provenancedataset standard

This study addresses the limited research credibility of existing synthetic electronic health records, which often suffer from inconsistencies between clinical processes and observations. The authors propose a two-stage integrated pipeline: first, a knowledge-guided generative model simulates high-fidelity patient trajectories by modeling nearly 32,000 clinical events; second, a large language model–based automated auditing module detects clinical contradictions, such as contraindicated medication prescriptions. Evaluated on 18,071 synthetic records, the method achieves high statistical fidelity (R² = 0.99), substantially reduces clinical inconsistencies, and enables downstream task performance comparable to or better than that achieved with real data—all without privacy leakage risks (F1 = 0.51). This work represents the first integration of knowledge-guided generation with LLM-driven clinical consistency auditing, significantly enhancing the clinical plausibility of synthetic medical records.

clinical consistencydata fidelityhealthcare data privacy

Hot Scholars

SZ

Shaoting Zhang

Shanghai AI Lab; SenseTime Research
Medical Image AnalysisComputer VisionFoundation Models
ZL

Zuozhu Liu

Assistant Professor, Zhejiang University/University of Illinois Urbana-Champaign
deep learningvision-language modelsmedical AI
WX

Weidi Xie

Shanghai Jiao Tong University | VGG, University of Oxford
Computer VisionAI for HealthcareAI for Science
SC

Shuheng Chen

University of Southern California
Machine LearningData SciencePredictive AnalyticsClinical Prediction