Score
Designs and applies frameworks, checklists, and automated or manual assessment methods to evaluate the accuracy, credibility, completeness, readability, and provenance of health-related information and content. Builds measurement instruments, rating protocols, and audits that generate quality scores, identify misinformation or bias, and produce actionable recommendations for improving health information quality.
This study systematically reviews NLP applications (2020–2024) for detecting, correcting, and mitigating medical misinformation—including clinical errors, false information, and LLM hallucinations. Following the PRISMA-ScR guidelines, it synthesizes BERT-based models, large language models (LLMs), rule-based systems, and hybrid approaches across tasks such as text classification, generative correction, and credibility scoring. Its primary contribution is a novel unified conceptual framework that integrates NLP strategies for all three problem types, alongside a cross-task methodology emphasizing clinical context modeling and explainable evaluation. The review identifies six core task-specific performance outcomes and critical bottlenecks: data privacy constraints, insufficient dynamic contextual modeling, and absence of standardized clinical validation metrics. Collectively, these findings provide both theoretical foundations and a practical roadmap for developing safe, transparent, and clinically deployable medical NLP systems.
Electronic health records (EHRs) are widely used in epidemiology and AI research, yet their data quality suffers from subgroup bias, systematic errors, and insufficient applicability assessment. To address these challenges, we propose a context-aware, automated medical data quality evaluation framework—the first to adapt software engineering principles of unit testing and coverage analysis to EHR validation. Our method integrates large language models (LLMs) for test case generation, medical knowledge anchoring, and research-context-driven data fitness analysis. Based on this, we develop MDPT, a tool comprising a test generator and executor. Evaluated on All of Us, MIMIC-III, and SyntheticMass datasets, MDPT generates 55–73 tests per cohort and detects 20–43 instances of data inconsistency or anomaly. The approach significantly improves both accuracy and interpretability in assessing EHR suitability for downstream research.
The absence of reliable, standardized quality assessment tools hinders clinical deployment of large language models (LLMs) for electronic health record (EHR) summarization. Method: We developed and validated PDSQI-9—the first standardized scale specifically designed to evaluate LLM-generated clinical summaries—based on multi-specialty real-world EHR data and outputs from GPT-4o, Mixtral, and Llama 3. The scale operationalizes a four-dimensional framework: organization, clarity, accuracy, and clinical utility. Rigorous validation included content, structural, criterion-related, discriminant, and generalizability validity assessments, conducted via semi-Delphi consensus, exploratory factor analysis, and multi-index reliability testing (Cronbach’s α = 0.879; ICC = 0.867). Results: PDSQI-9 demonstrates high psychometric validity and reliability, with statistically significant intergroup discrimination (p < 0.001). It provides a reproducible, generalizable benchmark for evaluating LLM-generated clinical summaries in real-world healthcare settings.
This study addresses the challenge of unreliable AI models and diminished clinical trust stemming from opaque data quality reporting in the secondary use of electronic health records (EHRs). To this end, the authors propose the first comprehensive framework for transparent data quality reporting across the entire EHR lifecycle. The framework innovatively distinguishes between data producers and consumers, explicitly defines five critical phases, and maps established data quality dimensions to specific workflow stages. Through iterative stakeholder and process analysis, a structured reporting mechanism is developed and validated on real-world datasets, demonstrating its ability to effectively trace the origins of data quality issues. The approach significantly enhances data interpretability, fitness-for-use, and governance efficacy, thereby providing a robust foundation for trustworthy AI development and clinical research.
Manual review of unstructured electronic health record (EHR) text to construct reference standards for large-scale database studies is time-consuming and labor-intensive. Method: We propose an NLP-driven, multi-wave adaptive sampling validation framework that integrates NLP-assisted annotation, quantitative bias analysis, and a predefined termination rule based on error convergence—dynamically optimizing both sample selection and stopping timing while preserving measurement accuracy. Results: Empirical evaluation shows that NLP reduces per-record review time by 40%; multi-wave sampling with termination criteria skips 77% of records requiring no manual review, with negligible impact (<0.5 percentage points) on final algorithm performance estimation bias. The framework significantly improves validation efficiency, feasibility, and scalability, offering a reproducible, resource-efficient, and standardized validation pathway for coded-algorithm-based health outcome measurement.
Existing AI-based clinical note evaluation methods suffer from misalignment between automated metrics and physician preferences, alongside the subjectivity and scalability limitations of expert reviews. Method: We propose Feedback2Checklist—a novel framework that distills large-scale, de-identified real-world clinical feedback into structured, interpretable, and actionable evaluation checklists, and builds an LLM-powered automated evaluator. Contribution/Results: Our approach significantly improves agreement between automated scores and physician preferences (+32.7% Spearman correlation), demonstrating high coverage, diversity, and robustness to quality degradation. Offline experiments show superior performance in identifying low-quality clinical notes compared to baselines, while ensuring strong clinical alignment and practical deployability.
Current medical research agents lack domain-specific evaluation mechanisms that rigorously assess scientific validity, methodological soundness, reproducibility, and boundary safety. This work proposes MedSkillAudit—the first skill auditing framework tailored for medical research agents—which employs a hierarchical, structured pipeline to evaluate skill readiness prior to deployment. The framework incorporates expert double-blind scoring (0–100), tiered release recommendations, and high-risk flags, and quantifies agreement between the system and human experts using ICC(2,1) and weighted Cohen’s kappa. Evaluated on 75 skills, the system achieved an ICC of 0.449, surpassing inter-human rater agreement (ICC = 0.300) and demonstrating closer alignment with consensus scores (SD = 9.5 vs. 12.4), thereby validating its effectiveness and reliability.
This study addresses the challenge of auditing the vast and complex content of Germany’s statutory health insurance websites, which exceeds manual review capacity and eludes detection by general-purpose AI tools due to nuanced medical, legal, and editorial issues. The authors propose a reproducible, multi-stage auditing framework that integrates deterministic rule-based filtering, large language model–assisted triage, temporal validity checks, and a dual-model comparison mechanism. Crucially, the approach distinguishes genuine content quality issues from superficial AI-generated signals without relying on AI-detection heuristics. Applied to 56,198 webpages, the framework prioritized content for review, producing 35,998 audit records and directing 21,452 pages into in-depth examination. In a stress test of 300 pages, it identified anomalies in 33.3% of cases, with the dual-model component achieving 75.8% agreement (Cohen’s κ = 0.532) across 182 matched instances.
Current health AI evaluation benchmarks lack standardized descriptions of user queries, limiting their ability to accurately reflect model applicability in real-world clinical settings. This study systematically identifies this “validity gap” and proposes adapting clinical trial reporting standards to create structured query profiles. Leveraging large language models, we automatically annotated 18,707 health-related queries from six public benchmarks using a 16-dimensional taxonomy capturing clinical context, topic, and intent. Our analysis reveals significant structural biases: existing benchmarks severely underrepresent complex diagnostic information such as laboratory tests, imaging, and raw clinical notes; safety-critical scenarios (e.g., self-harm) constitute less than 0.7% of queries; and coverage of pediatric, geriatric, and chronic disease populations is markedly insufficient—highlighting a substantial misalignment between current evaluation frameworks and actual clinical needs.