🤖 AI Summary
This study addresses the label efficiency bottleneck in surgical AI caused by the scarcity and high cost of expert annotations by proposing SURGE. The model is pretrained via self-supervised learning on over 30 million frames of surgical video, and its label efficiency is systematically evaluated across multiple benchmark tasks. The research reveals task-dependent label scaling laws, providing a theoretical blueprint for the efficient allocation of expert annotation resources in complex medical scenarios. Experimental results demonstrate that SURGE outperforms existing state-of-the-art methods across all tested benchmarks, surpassing even specialized models on complex reasoning tasks. Consequently, this work significantly enhances the few-shot generalization capability of surgical AI systems.
📝 Abstract
Developing label-efficient models is a central challenge in surgical AI due to the high cost and scarcity of expert annotation. While self-supervised foundation models adapt well to new tasks with minimal data, how label efficiency varies across different surgical tasks remains largely unexplored. Here, we introduce SURGE, a surgical foundation model trained on SurgSpectrum-30M+, the largest pretraining dataset comprising over 30 million frames, with checkpoints released to enable further research. We systematically evaluate label efficiency across 5 task categories and 15 benchmarks. These range from temporal and spatial scene understanding to fine-grained reasoning tied to instrument-anatomy interactions and safety-critical maneuvers. SURGE outperforms prior state-of-the-art on all benchmarks, even surpassing task-specific models on complex reasoning tasks. Crucially, we reveal a task-dependent scaling behavior: while scene understanding tasks saturate with minimal supervision, fine-grained reasoning tasks continue improving with substantially larger annotation budgets, providing a blueprint for allocating expert effort in complex domains. Code: https://github.com/CAMMA-public/SURGE