LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究针对急诊分诊中大型语言模型(LLMs)无法有效整合后续证据的问题,通过模拟和临床对话评估其在序列决策中的表现,发现LLMs依赖于主诉且随对话进展准确性下降。
📝 Abstract
Triage in the emergency department (ED) is a sequential decision process that unfolds turn by turn. Existing evaluations of large language models (LLMs) for triage use completed retrospective records and report performance close to that of physicians. We implement a methodology for evaluating LLMs on sequential triage, the task of predicting a triage acuity label from a growing prefix of a nurse-patient conversation. We evaluate six LLMs at five sequential checkpoints on two corpora: 425 LLM-generated (SIMULATED) and 50 physician-authored (CLINICIAN) conversations, both labelled under the Emergency Severity Index (ESI). Every model, measured by quadratic weighted kappa (QWK), degrades from moderate-to-substantial agreement on completed records to fair-to-moderate agreement at every sequential checkpoint. Controlled perturbations show that the label at every checkpoint is anchored on the chief complaint exchanges, and prompting interventions fail to lift this plateau. Models extract clinically relevant content from later turns, yet the surprisal of the true label rises across the checkpoints. So the model fails to integrate the evidence. Three expert clinicians on the same conversations reach a QWK of 0.887-0.929, while the best model reaches 0.295. Predictions concentrate at ESI-2 and ESI-3, and models agree with each other more than with the ground truth, so ensembling worsens the failure. Deploying LLMs for ED triage based on offline benchmarks alone misses this sequential failure.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Sequential Triage
Clinical Evidence Integration
Emergency Department
Decision Process
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sequential Triage
Large Language Models (LLMs)
Emergency Department (ED)
Quadratic Weighted Kappa (QWK)
Evidence Integration
🔎 Similar Papers
No similar papers found.
Dipankar Srirag
Dipankar Srirag
The University of New South Wales
Computational LinguisticsNatural Language ProcessingDialectalNLP
Haokai Zhao
Haokai Zhao
University of New South Wales
Deep Learning
A
Ashutosh Kumar
Independent Researcher
E
Eleanor Hopper
M
Michael Dalton
Q
Quoc Dung Nguyen
University of New South Wales, Sydney; Department of Rehabilitation Medicine, Campbelltown Hospital, Sydney
A
Aditya Joshi
University of New South Wales, Sydney
S
Salil S. Kanhere
University of New South Wales, Sydney
P
Padmanesan Narasimhan
University of New South Wales, Sydney