Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

๐Ÿ“… 2026-07-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses critical safety concerns regarding the use of large language models (LLMs) in autonomous clinical triage, particularly their inability to reliably identify โ€œcannot-missโ€ life-threatening conditions when presented with incomplete patient histories or high-risk scenarios. The work systematically uncovers a fundamental limitation rooted in the misalignment between LLMsโ€™ optimization objectives and clinical safety requirements: namely, their lack of sequential reasoning capabilities to actively gather information, broaden differential diagnoses, and escalate care under uncertainty. Integrating clinical reasoning frameworks, cognitive bias analysis, and behavioral evaluation of LLMs, the research demonstrates that while these models perform adequately in idealized settings, they exhibit subtle yet dangerous failures in real-world conditions characterized by incomplete data. Consequently, there is currently insufficient evidence to support the deployment of LLMs in unsupervised triage applications.
๐Ÿ“ Abstract
LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for symptom assessment and clinical decision support in diagnostic and treatment guidance, administrative documentation, and rules-based alert enhancement. This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop. For that task, the evidence of safety does not yet exist. The gap is not in medical knowledge but in the fidelity of clinical evaluation: a model optimized to continue the most probable text is not optimized to act safely when the safe answer is the improbable must-not-miss diagnosis. Safe triage is not the selection of the most likely diagnosis; it is a sequential decision under asymmetric cost, in which the single catastrophic miss outweighs many false alarms, and the decisive signal may be one the patient has not volunteered - and that the model has not been trained to seek. The core deficit is therefore one of information gathering under uncertainty. Under incomplete histories, LLM systems may fail to show the behaviors safe triage requires: broadening the differential; seeking the missing red flag; lowering the threshold for escalation; deferring judgement until sufficient information is obtained; and escalating concern where high-harm diagnoses remain unexcluded. These modes of failure for LLMs can be difficult to detect considering that evaluations to date often use complete, well-curated, confidence-gated simulations. The application of LLMs under these conditions may be amplified by assistant-like behaviors and positive bias, including credulity, agreeableness, and miscalibration - when these are not constrained by clinical triage logic.
Problem

Research questions and friction points this paper is trying to address.

clinical triage
large language models
diagnostic reasoning
patient safety
information gathering under uncertainty
Innovation

Methods, ideas, or system contributions that make the work stand out.

clinical reasoning
large language models
autonomous triage
information gathering under uncertainty
asymmetric decision-making
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
S
Shayndhan Sivanathan
Atman Labs; Oxford University Hospitals; Nuffield Department of Orthopaedics, Rheumatology and Musculoskeletal Sciences, University of Oxford; NIHR Oxford Biomedical Research Centre
S
Shravan Nageswaran
Atman Labs
M
Mehdi Zadem
Atman Labs
R
Ryaan Sultan
Atman Labs; Imperial College London
N
Nicolas von Mallinckrodt
Atman Labs; Technical University of Munich
M
Max Solovyev
Atman Labs
A
Alexey Matyushkin
Atman Labs
S
Sumon Sadhu
Atman Labs
G
Gabriele C DeLuca
Nuffield Department of Clinical Neurosciences, University of Oxford
S
Sanjeeva Jeyaretna
Department of Neurosurgery, Oxford; Nuffield Department of Surgical Sciences, University of Oxford
James Hillis
James Hillis
Facebook Reality Labs
biology and behavior
M
Manoj Ramachandran
Barts Health NHS Trust
Prakash Jayakumar
Prakash Jayakumar
Assistant Professor of Surgery and Perioperative Care
Health Care