Hallucination in Medical Imaging AI: A Cross-Modality Analytical Framework for Taxonomy, Detection, and Mitigation under Regulatory Constraints

📅 2026-06-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the critical issue of hallucinations in medical imaging AI—such as fabricated anatomical structures, left-right confusion, or erroneous measurements—which can lead to misdiagnosis and inappropriate treatment. The work proposes the first cross-modal hallucination taxonomy encompassing the entire imaging pipeline, systematically evaluating hallucination tendencies in both general-purpose and medical-specific foundation models. It integrates physical constraints, chain-of-thought prompting, and human-in-the-loop mechanisms to develop a detection and mitigation strategy aligned with FDA’s full lifecycle regulatory requirements. Key findings reveal that medical-specific models, due to overfitting, are paradoxically more prone to hallucinations; that combining multiple mitigation strategies effectively addresses diverse failure modes; and that radiologist review remains essential for safe clinical deployment.
📝 Abstract
AI systems are being deployed across medical imaging faster than their failure modes are understood. At this point in time, the failure of greatest clinical concern is hallucination: clinically plausible but factually incorrect outputs, including fabricated anatomical structures, missed findings, incorrect laterality, and invented measurements in generated reports, with direct consequences, for example, for biopsy decisions, staging, and treatment planning. This structured narrative synthesizes peer-reviewed studies, benchmark datasets, and FDA regulatory guidance across five imaging modalities to produce a cross-modality analysis of hallucination taxonomy, etiology, detection, and mitigation. Specifically, we address three questions in this study: (1) how can existing taxonomies be unified across modalities?, (2) how do medical-specialized foundation models hallucinate less than general-purpose ones?, and (3) which mitigation strategies are effective and compatible with FDA lifecycle oversight? We note that three taxonomic frameworks together cover the imaging pipeline in a way no single framework does alone. We also highlight that general-purpose foundation models outperform medical-specialized models on hallucination-specific benchmarks, indicating that narrow domain fine-tuning can introduce overfitting-induced confabulation. At the same time, the oversight of radiologists remains essential; for instance, a very high percentage of of AI-generated flags required expert correction before clinical use. Physics-informed architectural constraints, Chain-of-Thought prompting, and human-in-the-loop safeguards each address different failure modes and is effective when combined. All findings are mapped to the FDA's Total Product Lifecycle and Predetermined Change Control Plan frameworks, which treat hallucination management as a lifecycle obligation rather than a pre-deployment checklist.
Problem

Research questions and friction points this paper is trying to address.

hallucination
medical imaging AI
cross-modality
regulatory constraints
failure modes
Innovation

Methods, ideas, or system contributions that make the work stand out.

hallucination mitigation
cross-modality framework
foundation models
FDA regulatory compliance
physics-informed AI
🔎 Similar Papers
No similar papers found.
O
Omar Alshahrani
King Fahd University of Petroleum & Minerals, Saudi Arabia
M
Muzammil Behzad
King Fahd University of Petroleum & Minerals, Saudi Arabia; SDAIA-KFUPM Joint Research Center for Artificial Intelligence, Saudi Arabia