Clinician use of language models diverges from how the models are evaluated

πŸ“… 2026-10-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the critical misalignment between existing LLM clinical benchmarks and real-world workflows, which undermines the validity of deployment performance evaluations. Leveraging large-scale clinical query data, we integrate natural language processing, data mining, and expert validation to construct a comprehensive taxonomy and propose the RCQ-Map framework for systematically comparing public benchmarks against authentic usage distributions. Our quantitative analysis reveals that mainstream benchmarks cover only 31% of real-world task combinations and significantly underrepresent core clinical needs such as document processing. These findings demonstrate that current benchmark scores fail to effectively predict models’ actual clinical performance, providing essential evidence for developing future medical AI evaluation standards with higher ecological validity.
πŸ“ Abstract
Large language model (LLM) assistants are being deployed to clinicians across health systems, and judgments about their readiness rest largely on benchmark scores, most of them derived from examination questions or curated cases. A benchmark predicts performance in deployment only to the extent that its items resemble real use, yet whether benchmarks reflect the work these systems receive has rarely been measured. Here we analyze 127,833 queries sent by 6,342 physicians, advanced practice providers and nurses in 35 specialties to an institutional assistant during an eight-month roll-out. We characterize each query with RCQ-Map, a clinician-validated framework grounded in taxonomies of clinical questions and of LLM evaluation, which records its task, intent, answerability, missing information and potential harm. Documentation and administration (36.2%) and knowledge retrieval (28.9%) made up nearly two-thirds of use, and diagnosis 3.7%; more than a third of queries could not be answered well as posed. Applying RCQ-Map to 58 public benchmarks drawn from major evaluation suites and frontier model reports, which we assemble into the Clinical AI Benchmark Atlas, showed that the median benchmark contained no documentation requests and shared 31% of the task mix of real use, less than an even spread across task categories would. Benchmarks in suites designed to resemble clinical practice were individually no closer to real use than those used in frontier model reports. Benchmark scores therefore say little about how clinical AI performs on most of the work it is actually given, and evaluation should be matched to real clinical use.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Clinical AI Evaluation
Benchmark Alignment
Real-world Clinical Use
Innovation

Methods, ideas, or system contributions that make the work stand out.

RCQ-Map
Clinical AI Benchmark Atlas
Large Language Models
Benchmark Evaluation
Clinical Deployment
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
K
Krithik Vishwanath
Department of Neurosurgery, NYU Langone Health, New York, NY, USA
H
Haitong Lin
Department of Technology Management, NYU Tandon School of Engineering, New York University, New York, NY, USA
Anton Alyakin
Anton Alyakin
medical student at washington univesity
llmsneurosurgerynetworkscausality
J
Jin Vivian Lee
Department of Neurosurgery, NYU Langone Health, New York, NY, USA; Global AI Frontier Lab, New York University, New York, NY, USA
D
D. Brock Hewitt
Department of Surgery, NYU Langone Health, New York, NY, USA
J
Jie J. Yao
Department of Orthopedic Surgery, NYU Langone Health, New York, NY, USA
W
William Robert Small
Department of MCIT Health Informatics, NYU Langone Health, New York, NY, USA; Department of Medicine, NYU Langone Health, New York, NY, USA
H
Hammad A. Khan
Department of Neurosurgery, NYU Langone Health, New York, NY, USA
C
Cordelia Orillac
Department of Neurosurgery, NYU Langone Health, New York, NY, USA
A
Aakaash Varma
Division of Dermatology, Department of Medicine, NYU Langone Long Island, Mineola, NY, USA
B
Brandon Ye
Department of Neurosurgery, NYU Langone Health, New York, NY, USA; Johns Hopkins University School of Medicine, Baltimore, MD, USA
D
Daniel Alexander Alber
Department of Cardiothoracic Surgery, Stanford University School of Medicine, Stanford, CA, USA
Gustavo Stolovitzky
Gustavo Stolovitzky
Director, Biomed Data Sciences Hub, NYU Langone Health; Prof. Pathology Dep, NYU School of Medicine
Systems BiologyQuantitative BiologyGeneticsGenomicsData Science
B
Batia Wiesenfeld
Department of Management and Organizations, NYU Stern School of Business, New York, NY, USA
Oded Nov
Oded Nov
Professor, New York University
Human Computer InteractionHuman-AI InteractionDigital HealthAlgorithmic AdviceCSCW
W
Wei Wu
Faculty of Medicine, Macau University of Science and Technology, Taipa, Macao, China; Department of Big Data and Biomedical AI, College of Future Technology, Peking University and Peking-Tsinghua Center for Life Sciences, Beijing, China
Kang Zhang
Kang Zhang
Macau University of Science and Technology
ophthalmologygenomicsstem cellartificial intelligencetranslational medicine
Y
Yindalon Aphinyanaphongs
Department of MCIT Health Informatics, NYU Langone Health, New York, NY, USA; Department of Medicine, NYU Langone Health, New York, NY, USA; Department of Population Health, NYU Langone Health, New York, NY, USA
T
Tim Requarth
Department of Neuroscience, NYU Langone Health, New York, NY, USA
Eric Karl Oermann
Eric Karl Oermann
New York University
Artificial IntelligenceHuman Intelligence
T
The International Digital Twin Consortium in Healthcare
M
Medicine