Large language models exhibit unreliable updating of clinical judgment as patient evidence evolves

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of large language models (LLMs) to reliably update their judgments as clinical evidence evolves. We establish longitudinal belief updating as an independent evaluation dimension for LLM reliability. Through electronic health record trajectory analysis and controlled intervention experiments, we identify two failure modes: overreaction to deteriorating evidence and susceptibility to prior biases, demonstrating that conventional prompt engineering is ineffective against them. To address these limitations, we propose the Evidence-Validated Longitudinal Updating (EVLU) algorithm. Our findings indicate that although EVLU incurs a partial reduction in coverage, it effectively filters more reliable judgment revisions, thereby achieving an optimized trade-off between reliability and coverage.
📝 Abstract
Large language models (LLMs) are increasingly explored for clinical reasoning, but whether they appropriately revise judgments as patient evidence evolves remains unclear. We evaluated longitudinal belief updating using matched intensive-care trajectories from electronic health records. Across diverse LLMs, conditioning on a preceding judgment more often increased than reduced prediction error when estimates changed, replicated for a second endpoint. Controlled interventions revealed two failure modes. First, with preceding assessment fixed, models responded more strongly to worsening than matched improving respiratory evidence; this asymmetry persisted after headroom normalization at moderate and strong evidence levels. Second, with current evidence fixed, increasing prior risk from 10% to 90% shifted estimates by 26.2 percentage points, demonstrating causal influence of prior model beliefs. Prompting did not restore reliable updating. Evidence-Validated Longitudinal Update (EVLU) identified fewer, more reliable revisions, revealing a reliability-coverage trade-off. These findings establish longitudinal belief updating as a distinct dimension of LLM reliability.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Clinical Reasoning
Longitudinal Belief Updating
Prediction Error
LLM Reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Longitudinal Belief Updating
Clinical Reasoning
Evidence-Validated Longitudinal Update (EVLU)
Large Language Models
Reliability-Coverage Trade-off
🔎 Similar Papers
No similar papers found.
Min Zeng
Min Zeng
School of Computer Science and Engineering, Central South University
BioinformaticsMachine LearningDeep Learning
R
Rui Zhang
Division of Computational Health Sciences, Department of Surgery, University of Minnesota, Minneapolis, 55455, MN, USA.