RT-Super: Learning Tumor Segmentation from Longitudinal Images and Reports

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of tumor segmentation annotations by proposing a teacher-student knowledge distillation framework that eliminates the need for manual masks. The teacher network integrates longitudinal imaging with textual reports to generate high-quality pseudo-labels, while the student network employs a hybrid CNN-Transformer architecture to perform inference using only a single image. Furthermore, a novel loss function leveraging longitudinal consistency is introduced to optimize training. This approach overcomes the bottleneck of limited publicly available mask data, achieving highly accurate unsupervised segmentation across multiple tumor types, including esophageal cancer. Experimental results demonstrate that the proposed method outperforms existing mainstream public models in segmentation performance.
📝 Abstract
Multi-tumor segmentation is important for early cancer detection and allows radiologists to visualize, verify, and understand AI predictions. However, tumor segmentation masks are expensive, time-consuming, and unavailable for many tumor types in public data. Instead, hospitals have vast, readily available data that can guide segmentation: radiology reports, longitudinal images, and multi-phase images. We use this readily available data to substitute for tumor masks in training AI for tumor segmentation. To this end, we propose a new architecture, RT-Super. It has a teacher network, which analyzes the patient's longitudinal images and reports to create high-quality tumor masks. These masks train a student network, which sees a single image and no report. At inference, when longitudinal images and reports are unavailable, we use the student. RT-Super uses a new CNN-Transformer architecture and novel Consistency Losses that exploit tumor location consistency across longitudinal images. We train RT-Super to segment esophagus, uterus and spleen tumors, which have few or no public masks. Even without training masks, RT-Super can segment these tumors and surpass public AI models. Overall, we demonstrate that learning from longitudinal images, multi-phase images, and reports can overcome mask scarcity and advance multi-cancer detection and segmentation. Code: https://github.com/MrGiovanni/RT-Super
Problem

Research questions and friction points this paper is trying to address.

tumor segmentation
multi-tumor segmentation
mask scarcity
longitudinal images
radiology reports
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tumor Segmentation
Longitudinal Images
Knowledge Distillation
CNN-Transformer
Consistency Loss
P
Pedro R. A. S. Bassi
Johns Hopkins University
Wenxuan Li
Wenxuan Li
Johns Hopkins University
Imaging InformaticsComputer-aided Diagnosis
Hanxue Gu
Hanxue Gu
Duke University
Medical imagingDeep learningMachine learning
Jieneng Chen
Jieneng Chen
Johns Hopkins University
computer visionworld modelshealthrobotics
X
Xinze Zhou
Johns Hopkins University
Z
Zheren Zhu
University of California, San Francisco
S
Sezgin Er
University of Zurich
I
Ibrahim E. Hamamci
University of Zurich
B
Bjoern H. Menze
University of Zurich
G
Gulhan E. Akan
Istanbul Medipol University
K
Kang Wang
University of California, San Francisco
Yang Yang
Yang Yang
Associate Professor, University of California, San Francisco
Magnetic Resonance ImagingImage ReconstructionCardiovascular ImagingPerfusionAI
A
Alan L. Yuille
Johns Hopkins University
Zongwei Zhou
Zongwei Zhou
Assistant Research Professor, Johns Hopkins University
Medical Image AnalysisBiomedical InformaticsImaging InformaticsComputer-aided Diagnosis