KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limitation of existing medical AI benchmarks that evaluate only static diagnostic accuracy while neglecting dynamic clinical interactions such as patient interviewing and examination ordering. We introduce a clinical interaction benchmark comprising 333 physician-authored tasks, which evaluates models’ multi-turn dialogue, history-taking, and test-ordering capabilities through virtual patient simulations within an isolated sandbox environment. By incorporating authentic clinical workflows, decoupling inquiry from diagnostic scoring, and undergoing rigorous validation by over 35 clinical experts, the benchmark provides a comprehensive assessment framework. Experimental results reveal that although the best-performing model achieves 90.7% diagnostic accuracy, its success rate on complete clinical tasks remains below 30%, underscoring significant deficiencies in current large language models regarding dynamic clinical interaction capabilities.
πŸ“ Abstract
Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in different ways, and clinicians must obtain relevant history and determine which examinations are needed before reaching a diagnosis. Diagnostic accuracy alone therefore cannot establish whether an agent gathered essential information or conducted an appropriate clinical assessment. Furthermore, existing benchmarks lack professional clinicians' verification. To address this gap, we introduce KlinikeBench, a benchmark of 333 clinician-authored tasks, each providing an isolated sandbox environment with a virtual patient, clinical tools, and task-specific success criteria. More than 35 clinicians contributed to case authoring and benchmark evaluation. In an empirical study, clinicians gave simulated dialogues higher mean quality ratings than reference conversations, which is adapted from real conversation. In each task, an LM has a fixed budget of turns to communicate with the patient, ask about relevant history, request examinations, follow action constraints, and record a final diagnosis. We score these steps separately as well as together. Across 31 models and seven model families, the best-performing models (e.g., GPT-6-astra and Claude Opus 5) succeed on less than 30% of tasks, even though their diagnosis accuracy reaches 90.7%. Some models benefit from talking with the patient; others diagnose well from a complete chart but perform much worse in conversation. Overall, KlinikeBench provides a testbed for evaluating the full clinical encounter and reveals a substantial gap between diagnostic accuracy and performance in interactive clinical assessment.
Problem

Research questions and friction points this paper is trying to address.

clinical benchmark
language models
diagnostic accuracy
interactive clinical assessment
virtual patient
Innovation

Methods, ideas, or system contributions that make the work stand out.

Clinical Benchmark
Interactive Assessment
Virtual Patient Sandbox
Language Model Evaluation
Diagnostic Accuracy
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
X
Xueting Fang
Zhejiang University
Zehui Li
Zehui Li
PhD, Imperial College London
Machine LearningDeep LearningBioinformatics
Y
Yang Yang
Nanchang University
C
Camilla Giovino
University of Toronto
S
Shubh K. Patel
University of Toronto
S
Shailly Prajapati
University of Toronto
Vallijah Subasri
Vallijah Subasri
University Health Network
responsible AIprecision medicinecomputational biologymachine learning for healthcare
Caihua Shan
Caihua Shan
Microsoft Research Asia
Graph LearningAI for ScienceCrowdsoucing