🤖 AI Summary
This study addresses the scarcity of tumor segmentation annotations by proposing a teacher-student knowledge distillation framework that eliminates the need for manual masks. The teacher network integrates longitudinal imaging with textual reports to generate high-quality pseudo-labels, while the student network employs a hybrid CNN-Transformer architecture to perform inference using only a single image. Furthermore, a novel loss function leveraging longitudinal consistency is introduced to optimize training. This approach overcomes the bottleneck of limited publicly available mask data, achieving highly accurate unsupervised segmentation across multiple tumor types, including esophageal cancer. Experimental results demonstrate that the proposed method outperforms existing mainstream public models in segmentation performance.
📝 Abstract
Multi-tumor segmentation is important for early cancer detection and allows radiologists to visualize, verify, and understand AI predictions. However, tumor segmentation masks are expensive, time-consuming, and unavailable for many tumor types in public data. Instead, hospitals have vast, readily available data that can guide segmentation: radiology reports, longitudinal images, and multi-phase images. We use this readily available data to substitute for tumor masks in training AI for tumor segmentation. To this end, we propose a new architecture, RT-Super. It has a teacher network, which analyzes the patient's longitudinal images and reports to create high-quality tumor masks. These masks train a student network, which sees a single image and no report. At inference, when longitudinal images and reports are unavailable, we use the student. RT-Super uses a new CNN-Transformer architecture and novel Consistency Losses that exploit tumor location consistency across longitudinal images. We train RT-Super to segment esophagus, uterus and spleen tumors, which have few or no public masks. Even without training masks, RT-Super can segment these tumors and surpass public AI models. Overall, we demonstrate that learning from longitudinal images, multi-phase images, and reports can overcome mask scarcity and advance multi-cancer detection and segmentation. Code: https://github.com/MrGiovanni/RT-Super