INFORM-CT: INtegrating LLMs and VLMs FOR Incidental Findings Management in Abdominal CT

๐Ÿ“… 2025-12-10
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Current clinical detection and reporting of incidental findings (IFs) in abdominal CT rely heavily on manual interpretation, resulting in low efficiency and poor inter-observer consistency. To address this, we propose a planning-execution intelligent agent framework that synergistically integrates large language models (LLMs) and vision-language models (VLMs). Guided by clinical guidelines, the LLM performs interpretable, programmatic reasoning to generate task plans, which orchestrate VLMs and organ-specific segmentation models for IF detection, classification, and structured report generationโ€”fully supported by automated Python script compilation and execution. This work introduces, for the first time in medical imaging analysis, an LLM-driven end-to-end task planning paradigm with coordinated multimodal model execution, significantly enhancing guideline adherence and reasoning transparency. Evaluated on a three-organ abdominal CT benchmark, our method achieves superior end-to-end accuracy and processing efficiency compared to state-of-the-art VLM-only approaches.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Large Vision ModelsIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Agentic search
๐Ÿ“ Abstract
Incidental findings in CT scans, though often benign, can have significant clinical implications and should be reported following established guidelines. Traditional manual inspection by radiologists is time-consuming and variable. This paper proposes a novel framework that leverages large language models (LLMs) and foundational vision-language models (VLMs) in a plan-and-execute agentic approach to improve the efficiency and precision of incidental findings detection, classification, and reporting for abdominal CT scans. Given medical guidelines for abdominal organs, the process of managing incidental findings is automated through a planner-executor framework. The planner, based on LLM, generates Python scripts using predefined base functions, while the executor runs these scripts to perform the necessary checks and detections, via VLMs, segmentation models, and image processing subroutines. We demonstrate the effectiveness of our approach through experiments on a CT abdominal benchmark for three organs, in a fully automatic end-to-end manner. Our results show that the proposed framework outperforms existing pure VLM-based approaches in terms of accuracy and efficiency.
Problem

Research questions and friction points this paper is trying to address.

Automates incidental findings management in abdominal CT scans
Integrates LLMs and VLMs for detection and reporting efficiency
Improves accuracy over pure VLM-based methods via planner-executor framework
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-based planner generates Python scripts for automation
Executor uses VLMs and segmentation for medical image analysis
Plan-and-execute agentic framework improves incidental findings management
I
Idan Tankel
GE Healthcare Technology and Innovation Center, Niskayuna, NY, USA
N
Nir Mazor
GE Healthcare Technology and Innovation Center, Niskayuna, NY, USA
R
Rafi Brada
GE Healthcare Technology and Innovation Center, Niskayuna, NY, USA
C
Christina LeBedis
Boston Medical Center, Boston, MA, USA
Guy Ben-Yosef
Guy Ben-Yosef
Principal Scientist at GE Research
Computer VisionMachine LearningCognitive ScienceMedical Imaging