🤖 AI Summary
This study addresses the weak alignment between clinical descriptions and images in conventional colonoscopy reports, which typically lack frame-level annotations. The authors present EndoCLIP, a colonoscopy-specific vision–language foundation model trained on 125,756 lesion-level image–text pairs automatically extracted from 280,000 unstructured procedure-level reports—the first such large-scale dataset derived directly from routine clinical documentation. By transforming standard clinical narratives into scalable supervisory signals, EndoCLIP enables task specification via natural language. Evaluated under zero-shot and linear probing settings, EndoCLIP outperforms both general-purpose and biomedical vision–language models across lesion-level image–text retrieval, structured report generation, and six multicenter classification tasks. Notably, its performance in benign–malignant lesion classification under linear probing approaches that of a panel of twelve expert endoscopists.
📝 Abstract
Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images. Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonoscopy records. Across lesion-level image-text retrieval, structured report generation and six multi-centre clinical classification tasks, EndoCLIP outperforms general-purpose and biomedical vision-language encoders in both zero-shot and linear-probe settings. On benign-versus-malignant classification, its linear probe approaches the performance of expert readers in a blinded study involving 12 endoscopists. These results suggest that recovering finding-to-frame correspondence can transform routine documentation into scalable supervision, enabling clinical targets to be specified in language rather than separately annotated for each task.