Score
Designing and executing processes to recruit, validate, and engage clinical experts for studies, including collecting and validating physician annotations, rating clinical importance, and running blinded, specialty-matched evaluations.
This work addresses the time-consuming and cognitively demanding process of drafting clinical trial eligibility criteria, which existing automated approaches often fail to support effectively due to reliance on structured inputs or limited controllability. The authors propose a guided generation framework grounded in semantic axes—such as demographics, laboratory values, and behavioral factors—that enables clinicians to steer large language models toward generating appropriate eligibility criteria without specifying exact entities. By incorporating an intermediate control mechanism and a reusable, multidimensional scoring system, the method strikes a balance between controllability and clinical utility. Experimental results demonstrate that the approach significantly outperforms unguided baselines across automatic metrics, standardized scoring rubrics, and clinician evaluations, thereby enhancing both the interpretability and practical value of AI-assisted clinical trial design.
Clinical trial eligibility matching is critical for ensuring scientific rigor and patient safety, yet conventional manual screening suffers from low efficiency and high error rates. This study presents a systematic review of NLP-driven matching approaches published between 2015 and 2024. We propose a novel paradigm integrating rule-based engines, named entity recognition (NER), contextual embeddings (e.g., BERT), and ontology-based normalization using UMLS and SNOMED CT—unifying explainable AI with standardized medical ontologies to enhance model transparency and trustworthiness. Empirical evaluation demonstrates substantial improvements in both matching accuracy and processing speed. Furthermore, we identify three persistent challenges: data fragmentation across sources, inconsistent annotation practices, and limited generalizability across clinical sites. Finally, we outline a new research direction toward joint semantic-temporal modeling to better capture dynamic eligibility criteria.
Clinical workflow fragmentation severely impedes efficiency: heterogeneous scripting, ad-hoc model ensembles, and lack of data-driven modality identification and standardized outputs result in high deployment overhead, costly monitoring, and poor interoperability. To address this, we propose a healthcare-first vision-language unified framework that pioneers the use of a single vision-language model (VLM) for two-tier clinical decision-making—first, an auditable, three-stage routing mechanism matches inputs to expert-defined model cards; second, domain-specific multi-task joint inference (with early-exit capability and candidate arbitration) adheres to clinical risk constraints. Leveraging phased prompting, a candidate answer selector, and specialty-specific fine-tuning, our framework unifies modality identification, abnormality classification, model selection, and multi-task reasoning. Evaluated across gastroenterology, hematology, ophthalmology, and pathology, our single-model solution achieves performance on par with specialized models while substantially reducing deployment complexity, operational overhead, and integration effort.
This study addresses the low efficiency and geographical constraints of traditional clinical trial recruitment by proposing, for the first time, a large language model (LLM)-based paradigm to identify and assess potential participants from social media. We introduce TRIALQA, the first annotated dataset designed for real-world trial eligibility determination, comprising colon and prostate cancer–related social media texts and supporting multi-hop reasoning and participant motivation identification. We systematically evaluate six training and inference strategies across seven state-of-the-art LLMs on eligibility criterion matching and willingness-to-participate classification, with rigorous multi-round human annotation ensuring high data quality. Results demonstrate that LLMs exhibit promising capability in comprehending complex medical criteria but remain limited in multi-step logical reasoning and fine-grained judgment. This work establishes a new benchmark, a novel publicly available dataset, and a reproducible methodology for AI-driven precision clinical trial recruitment.
Methodological gaps persist in leveraging expert opinion for borrowing treatment effects or parameters in clinical trials for both standard and rare diseases. Method: We conducted a systematic literature review—including database searching and citation tracking—to identify 41 relevant studies, and performed a novel mapping and classification of expert elicitation and aggregation approaches used in trial design and analysis. Contribution/Results: We identified six elicitation strategies and ten aggregation methods, revealing that existing techniques are broadly applicable but lack contextual adaptation to clinical trial decision-making and formal methodological frameworks. This study introduces the first clinical trial–oriented methodology map for expert opinion integration, characterizing key practical features—including expert sample size, training protocols, parameter types, and distributional assumptions. The map informs the development of context-sensitive, empirically verifiable methods for integrating expert knowledge into clinical trial design and analysis.
This study addresses the gap between evidence-based medicine and personalized clinical practice by examining whether physicians’ experience-driven treatment decisions outperform the average optimal therapy recommended by randomized controlled trials (RCTs), which often overlook individual heterogeneity. By integrating RCT and observational cohort data from the same population, the authors propose a “gain score” to quantify the benefit of physician-assigned strategies relative to trial-based recommendations. Within a nested study design, they derive sharp nonparametric bounds on the proportion of patients for whom physician strategies are superior—a result established for the first time without parametric assumptions. Leveraging causal inference and partial identification techniques, this work provides a data-driven framework to determine when clinical discretion should supersede population-average guidelines, thereby bridging the divide between standardized evidence and individualized care.
This study introduces the novel concept of “clinical trial engineering”—the systematic manipulation of statistical analyses to generate misleading clinical trial evidence in support of drug approval, distinct from conventional paper mills. Focusing on 23 studies linked to Iran’s CinnaGen and its subsidiary Orchid Pharmed, the authors applied the INSPECT-SR credibility framework, integrating PubMed literature screening, raw data verification, and co-authorship network analysis to systematically evaluate evidentiary reliability. The investigation uncovered 180 issues spanning nine categories of systemic bias, including incomplete reporting, arithmetic errors, and design flaws. These findings reveal a structural pattern of research manipulation driven by commercial pressures, publication incentives, and permissive regulatory pathways, prompting regulatory agencies to reassess the credibility of the associated clinical evidence.
This study addresses the limitations of existing medical agent evaluation benchmarks, which struggle to simulate the long-horizon, multi-step, and verifiable workflows characteristic of real-world clinical practice. To bridge this gap, the authors introduce a novel benchmark built upon a real electronic health record (EHR) system, comprising 100 cross-specialty tasks derived from actual consultation cases spanning 21 specialties and diverse clinical processes. Agents are required to perform integrated operations including data retrieval, clinical reasoning, system interaction, and documentation. For the first time, executable and verifiable tasks are deployed via a standard EHR API, complemented by a structured checkpoint mechanism enabling fine-grained assessment. Experimental results reveal that even the best-performing among 13 leading large language model agents achieves only a 46% one-shot success rate, with open-source models reaching just 19%, underscoring a substantial gap between current capabilities and real-world clinical demands.