Score
Designs and executes validation frameworks, test plans, and evidence packages to demonstrate that clinical predictive, diagnostic, or decision‑support models meet safety, performance, fairness, robustness, and reliability requirements; performs statistical performance analyses, calibration and error analysis, subgroup and bias assessments, and prospective or retrospective validation studies. Prepares and maintains regulatory documentation and risk-management artifacts—requirements traceability, clinical evaluation reports, change‑control and post‑market surveillance plans—and evaluates model changes for continued compliance with applicable clinical regulatory standards.
This study addresses the need for more reliable regulatory decision-making in Bayesian clinical trials by systematically calibrating Bayesian success criteria to control decision errors. It establishes the first theoretical correspondence between Bayesian decision error metrics and frequentist operating characteristics—specifically Type I and Type II error rates—and proposes a practical calibration strategy grounded in this relationship. The approach is illustrated through a case study on a revascularization trial in cardiogenic shock. To facilitate adoption under the FDA’s emerging Bayesian framework, the authors also developed an interactive Shiny web application that enables sponsors and regulators to efficiently and reliably formulate decisions while maintaining rigorous error control.
This study addresses the frequent violations of clinical coding standards—such as ICD-10, CPT, and HL7 FHIR—by large language models when generating structured medical data, which impedes integration with electronic health record systems. To mitigate this, the authors propose and validate a closed-loop verification-and-repair framework that automatically detects and iteratively corrects formatting errors. The approach is evaluated using three open-source models—Qwen2.5-7B, Llama3.1-8B, and Gemma2-9B—deployed locally across 320 clinical scenarios. Results demonstrate a substantial improvement in schema compliance across all models, achieving an overall adherence rate of 99.0% and increasing individual model performance by 7.8 to 12.5 percentage points. Notably, 96% of detected errors were attributable to repairable representation-layer issues, with most resolved within one or two correction rounds, effectively compensating for the models’ limited understanding of healthcare IT standards.
This study addresses the challenge of insufficient sample sizes in clinical prediction model development, which often leads to overfitting and poor generalizability, while existing methods struggle to reliably estimate the minimum required sample size. The authors propose a simulation-based, safeguard-oriented sample size calculation framework that innovatively integrates learning curves, Gaussian process optimization, and user-defined performance criteria, enabling a model-agnostic design. Implemented in the open-source R package pmsims, this approach efficiently accommodates diverse modeling techniques and evaluation metrics while explicitly quantifying performance uncertainty. Case studies demonstrate that pmsims offers marked advantages over current tools in terms of flexibility, computational efficiency, and broad applicability.
Clinical risk-scoring benchmarks such as MedCalc-Bench suffer from the entrenchment of historically erroneous models as de facto “gold standards,” particularly problematic when used as reward signals in reinforcement learning, where such biases are amplified. Method: We propose a “living documentation” benchmark paradigm featuring physician-in-the-loop, low-burden dynamic curation: GRPO-based RL fine-tuning of Qwen3-8B, multi-stage agentic verification, LLM-driven logical consistency checking, and clinical-knowledge-guided controversy identification—enabling scalable, trustworthy re-annotation. Contribution/Results: Experiments uncover substantial label noise; post-correction, model accuracy improves by 8.7 percentage points. This work is the first systematic demonstration that safety-critical domain benchmarks require continuous governance, establishing a novel human–AI collaborative paradigm for dynamic benchmark auditing.
This study addresses critical safety concerns in the clinical deployment of large language models (LLMs), particularly their inability to reliably screen for high-risk conditions and provide verifiable diagnostic reasoning. To this end, the authors propose AegisDx, a novel framework that introduces, for the first time, a safety-prioritized hypothetico-deductive reasoning mechanism. AegisDx integrates role-specialized LLM components, structured intermediate outputs, a medical evidence retrieval interface, and a validation gating mechanism to jointly ensure systematic exclusion of life-threatening diseases and enable traceable differential diagnosis. Experimental results demonstrate that AegisDx significantly outperforms baseline models in top-3 diagnostic accuracy across multiple medical case collections from peer-reviewed journals, achieves a physician-blind safety rating of 4.55 out of 5, and improves high-risk disease detection by 26 percentage points.
Current approaches to sample size calculation for clinical prediction models typically neglect the impact of missing data, often resulting in overfitting and poor calibration. This study is the first to integrate missing data mechanisms and handling strategies—such as multiple imputation—into a posterior-distribution-based sample size framework. Through simulation studies and Expected Value of Perfect Information (EVPI) analyses, the research quantifies how missingness affects model performance. Findings reveal that under common missing data scenarios, even when existing minimum sample size criteria are met, calibration slopes frequently fall below 0.9. In certain settings, nearly twice the conventional sample size is required to achieve performance comparable to that with complete data, underscoring both the necessity and feasibility of dynamically adjusting sample size requirements in the presence of missing data.
This study addresses the challenge of error-prone manual verification of tables, figures, and listings (TFLs) in clinical trial reports, which often fails to detect structural or logical inconsistencies. The authors propose PROVE, a novel framework that leverages large language models (LLMs) for semantic parsing and evidence tracing of TFL content, integrated with a programmable rule engine to perform deterministic numerical and logical validation against SDTM/ADaM standards. Designed as a multi-agent architecture, PROVE combines LLM-driven semantic understanding with rule-based checks to enable auditable, configurable automated cross-verification. The approach achieves 100% accuracy under exact label matching; when confronted with linguistic variations, LLM assistance boosts recall from 0.588 to 0.993 and F1 score from 0.735 to 0.996.
This study addresses the challenge of safely and effectively reducing sample sizes in registered clinical trials while adhering to regulatory requirements. Building upon the FDA’s seven-step risk assessment framework, the work presents the first systematic application of AI model trustworthiness evaluation guidelines to the context of sample size reduction. By constructing prognostic covariates, conducting risk-informed model development and validation, and integrating these with statistical re-estimation methods, the approach recalculates the required trial sample size. Demonstrated in a randomized controlled trial for Alzheimer’s disease, the methodology enabled prospective sample size reduction, substantially shortening trial duration, lowering costs, and accelerating the availability of effective therapies. This provides a generalizable, AI-driven framework for enhancing the efficiency of drug development.
Legacy clinical reporting systems hinder AI integration and impede drug development and pharmacovigilance due to opaque outputs and the absence of machine-readable intermediate representations. This work proposes a non-intrusive, metadata-driven framework that bridges legacy components—without modifying their source code—through mapping layers, a typed intermediate representation (IR), and coordinator wrappers, thereby transforming their outputs into structured data suitable for large language models (LLMs) and enabling progressive replacement. Validated on 558 SAS components (373k lines of code), the approach achieves immediate AI readiness in coexistence mode, reduces proprietary code by 92% after optional integration, and demonstrates unit-level consistency exceeding 80% across 11 of 14 report types (mean: 82.7%, peak: 99.2%). Five reports achieved 100% compliance on the CDISCPilot01 benchmark, and the framework successfully enabled LLM-driven automated pharmacovigilance, table summarization, and trial configuration generation.