๐ค AI Summary
Current clinical detection and reporting of incidental findings (IFs) in abdominal CT rely heavily on manual interpretation, resulting in low efficiency and poor inter-observer consistency. To address this, we propose a planning-execution intelligent agent framework that synergistically integrates large language models (LLMs) and vision-language models (VLMs). Guided by clinical guidelines, the LLM performs interpretable, programmatic reasoning to generate task plans, which orchestrate VLMs and organ-specific segmentation models for IF detection, classification, and structured report generationโfully supported by automated Python script compilation and execution. This work introduces, for the first time in medical imaging analysis, an LLM-driven end-to-end task planning paradigm with coordinated multimodal model execution, significantly enhancing guideline adherence and reasoning transparency. Evaluated on a three-organ abdominal CT benchmark, our method achieves superior end-to-end accuracy and processing efficiency compared to state-of-the-art VLM-only approaches.
๐ Abstract
Incidental findings in CT scans, though often benign, can have significant clinical implications and should be reported following established guidelines. Traditional manual inspection by radiologists is time-consuming and variable. This paper proposes a novel framework that leverages large language models (LLMs) and foundational vision-language models (VLMs) in a plan-and-execute agentic approach to improve the efficiency and precision of incidental findings detection, classification, and reporting for abdominal CT scans. Given medical guidelines for abdominal organs, the process of managing incidental findings is automated through a planner-executor framework. The planner, based on LLM, generates Python scripts using predefined base functions, while the executor runs these scripts to perform the necessary checks and detections, via VLMs, segmentation models, and image processing subroutines.
We demonstrate the effectiveness of our approach through experiments on a CT abdominal benchmark for three organs, in a fully automatic end-to-end manner. Our results show that the proposed framework outperforms existing pure VLM-based approaches in terms of accuracy and efficiency.