CoralBay: A Self-Supervised CT Foundation Model

📅 2026-06-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing two-dimensional self-supervised methods struggle to effectively model the spatial continuity, anatomical structures, and Hounsfield Unit representations inherent in three-dimensional CT imaging, thereby limiting their transfer performance in medical tasks. To address this, this work proposes CoralBay—the first self-supervised foundation model tailored for 3D CT scans—built upon a hierarchical 3D Swin Transformer architecture that integrates multi-scale feature concatenation with an extended self-distillation mechanism (a DINO variant) to jointly learn global semantics and local details. We introduce self-distillation to 3D medical imaging for the first time and establish the first unified, reproducible benchmark for 3D radiology evaluation. Experimental results demonstrate that CoralBay significantly outperforms existing approaches across diverse downstream tasks, advancing standardized evaluation in 3D medical visual representation learning.
📝 Abstract
Self-supervised learning has enabled large-scale pre-training on 2D natural images, producing general-purpose visual representations that transfer effectively across tasks. However, many medical imaging modalities, such as CT scans, are inherently three-dimensional and differ fundamentally from natural images in both structure and semantics. Volumetric modalities capture spatial continuity, organ anatomy, and intensity-based tissue properties (e.g., Hounsfield Units), which are not adequately modeled by 2D pre-training. To bridge this gap, we introduce CoralBay, a self-distillation framework that extends DINO by using a hierarchical 3D Swin backbone and applying self-distillation to concatenated multi-scale features, enabling data-efficient self-supervised learning of rich spatial representations that encode both global semantics and fine-grained local structure. As a result, CoralBay transfers effectively to a wide range of downstream radiological tasks, demonstrating strong and consistent performance across diverse anatomical targets. In addition, we contribute to the open-source \eva framework by introducing a public, reproducible 3D radiology leaderboard that unifies multiple datasets and establishes a standardized benchmark for evaluating volumetric representation learning methods.
Problem

Research questions and friction points this paper is trying to address.

self-supervised learning
3D medical imaging
CT scans
volumetric representation
transfer learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-supervised learning
3D medical imaging
self-distillation
Swin Transformer
foundation model
💼 Related Jobs
No related jobs found.
I
Ioannis Gatopoulos
kaiko.ai
N
Nicolas Känzig
kaiko.ai
Sebastian Otálora
Sebastian Otálora
MLE @ Kaiko.ai
machine learningbiomedical image analysis
F
Fei Tang
kaiko.ai