Incidental information contaminates patient notes and disrupts clinical reasoning in large language models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of large language models to contamination by irrelevant speech during clinical documentation and reasoning. Through large-scale dialogue simulation, transcription analysis, and open-source model evaluation, we systematically assess the sensitivity of frontier models to chitchat and ambient speech. Our quantitative findings reveal a 35% note insertion rate and a 48.2% transcription leakage rate, exposing substantial fragility in clinical reasoning pipelines. We further propose a “dual-coding” hypothesis that elucidates the interference mechanism whereby irrelevant information and clinical reasoning compete for shared neural components. This work establishes a new paradigm requiring rigorous anti-contamination evaluation prior to clinical deployment, providing critical empirical evidence for constructing robust safety guardrails in healthcare AI systems.
📝 Abstract
Large language models (LLMs) are increasingly relied upon to support ambient documentation and clinical reasoning. Here we examine the impact of a failure mode shared between these two applications by assessing their sensitivity to information incidental to the patient encounter. In 576 patient-clinician dialogues, we found that frontier models inserted small-talk exchanges into 35% of notes, while mean quality scores changed by at most 0.20 points on five-point scales. In 3.7% of frontier notes, models misattributed the asides or used them clinically. In 57 mock recorded consultations, background speech from a separate patient encounter at -10 dB leaked into 48.2% of transcripts, with contamination detected in 5.3% of downstream notes generated by four open-weight models. We propose a dual encoding hypothesis of clinical reasoning and distraction in LLMs, with preliminary evidence that LLM components associated with disruption by incidental information also support clinical reasoning. These findings support evaluating resistance to incidental information before clinical use, with safeguards that prevent contamination while preserving clinical reasoning.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Clinical Reasoning
Incidental Information
Ambient Documentation
Information Contamination
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Clinical Reasoning
Incidental Information
Dual Encoding Hypothesis
Ambient Documentation
🔎 Similar Papers
No similar papers found.
K
Krithik Vishwanath
Department of Neurosurgery, NYU Langone Health, New York, NY, USA
B
Brandon Ye
Johns Hopkins University School of Medicine, Baltimore, MD, USA
Anton Alyakin
Anton Alyakin
medical student at washington univesity
llmsneurosurgerynetworkscausality
J
John E. Markert
University of Alabama at Birmingham Heersink School of Medicine, Birmingham, AL, USA
A
Aaron Hsieh
Perelman School of Medicine at the University of Pennsylvania, Philadelphia, PA, USA
M
Michał Mańkowski
Department of Surgery, NYU Langone Health, New York, NY, USA
E
Eric K. Oermann
Department of Neurosurgery, NYU Langone Health, New York, NY, USA; Global AI Frontier Lab, New York University, New York, NY, USA; Department of Radiology, NYU Langone Health, New York, NY, USA; Neuroscience Institute, NYU Langone Health, New York, NY, USA; Center for Data Science, New York University, New York, NY, USA