🤖 AI Summary
This study addresses the scarcity and high annotation cost of training data for 3D visual grounding in abdominal CT by, for the first time, leveraging 2D annotations from routine clinical PACS as implicit supervision signals. Through metadata matching and promptable segmentation models, we construct an automated data pipeline that generates 105,000 phrase-mask-volume triplets without any additional manual annotation. Based on this pipeline, we introduce the LocusBench benchmark and propose the LocusCT model. In two evaluation settings, our approach achieves hit rates of 0.725 and 0.773, respectively, significantly outperforming existing methods. This work substantially lowers the barrier to constructing large-scale datasets for 3D visual grounding in medical imaging.
📝 Abstract
Visual grounding models can enhance radiology workflows by linking report findings to image regions. This is particularly valuable for 3D CT, where findings often occupy a tiny fraction of the volume. Training 3D grounding models requires large sets of paired phrases and regions, and building such datasets is expensive, requiring radiologists to annotate images by hand. We posit that this supervision is already created implicitly during routine reporting, as radiologists frequently place 2D annotations (e.g., distance measurement, arrows) on key images to make measurements and to support report interpretation. We introduce an automated pipeline that converts these routine clinical annotations into large-scale phrase-region supervision for 3D visual grounding. The pipeline links each annotation to the corresponding finding in the report through metadata matching, then uses a promptable 3D segmentation model to convert the 2D annotation into a volumetric mask. This produces phrase-mask-volume datasets without requiring additional radiologist annotation. Applied to a single institution's clinical picture archiving and communication system (PACS), our approach generated 105K phrase-mask-volume triplets from 59K abdominal CT exams. We also introduce two abdominal CT grounding benchmarks, LocusBench-Onc and LocusBench-ED, which comprise 240 oncology and 260 emergency-department radiologist-reviewed phrase-mask-volume triplets, respectively, with the latter spanning 13 distinct categories such as appendicitis, hematoma, and hernia. We further introduce LocusCT, a 3D visual grounding model trained on this dataset, which achieves hit rates of 0.725 on LocusBench-Onc and 0.773 on LocusBench-ED, substantially outperforming comparator models. These results show that routine PACS annotations are a scalable, previously unused source of supervision for 3D visual grounding.