A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the weak alignment between clinical descriptions and images in conventional colonoscopy reports, which typically lack frame-level annotations. The authors present EndoCLIP, a colonoscopy-specific vision–language foundation model trained on 125,756 lesion-level image–text pairs automatically extracted from 280,000 unstructured procedure-level reports—the first such large-scale dataset derived directly from routine clinical documentation. By transforming standard clinical narratives into scalable supervisory signals, EndoCLIP enables task specification via natural language. Evaluated under zero-shot and linear probing settings, EndoCLIP outperforms both general-purpose and biomedical vision–language models across lesion-level image–text retrieval, structured report generation, and six multicenter classification tasks. Notably, its performance in benign–malignant lesion classification under linear probing approaches that of a panel of twelve expert endoscopists.
📝 Abstract
Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images. Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonoscopy records. Across lesion-level image-text retrieval, structured report generation and six multi-centre clinical classification tasks, EndoCLIP outperforms general-purpose and biomedical vision-language encoders in both zero-shot and linear-probe settings. On benign-versus-malignant classification, its linear probe approaches the performance of expert readers in a blinded study involving 12 endoscopists. These results suggest that recovering finding-to-frame correspondence can transform routine documentation into scalable supervision, enabling clinical targets to be specified in language rather than separately annotated for each task.
Problem

Research questions and friction points this paper is trying to address.

colonoscopy
vision-language models
clinical documentation
image-text alignment
lesion annotation
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language model
colonoscopy
lesion-level alignment
foundation model
zero-shot learning
🔎 Similar Papers
No similar papers found.
Jia Yu
Jia Yu
Co-founder, Wherobots Inc.; Assistant Professor of Computer Science, Washington State University
Database systemsData managementGeospatial databasesGIS
Y
Yan Zhu
Shanghai Collaborative Innovation Center of Endoscopy, Shanghai, China.
Y
Yili He
Digital Medical Research Center, School of Basic Medical Sciences, Fudan University, Shanghai, China.
Z
Zilong Wang
Microsoft Research Asia, Shanghai, China.
Xinyang Jiang
Xinyang Jiang
Microsoft Research Asia
Computer VisionReIDDeep Learning
P
Peiyao Fu
Shanghai Collaborative Innovation Center of Endoscopy, Shanghai, China.
R
Ruijie Yang
Digital Medical Research Center, School of Basic Medical Sciences, Fudan University, Shanghai, China.
T
Tianyi Chen
Shanghai Collaborative Innovation Center of Endoscopy, Shanghai, China.
Siyuan Li
Siyuan Li
Shanghai Jiao Tong University
Trustworthy LLM AgentsEdge Intelligence
Zhihua Wang
Zhihua Wang
City University of Hong Kong
Computer VisionBiomedical EngineeringRobotics
Fei Wu
Fei Wu
Professor of Computer Science, Zhejiang University
Multimedia RetrievalSparse RepresentationMachine LearningKnowledge Graph
Q
Quanlin Li
Endoscopy Centre and Endoscopy Research Institute, Zhongshan Hospital, Fudan University, Shanghai, China.
Xian Yang
Xian Yang
University of Manchester
Artificial IntelligenceMachine LearningHealthcare AINatural Language Processing
P
Pinghong Zhou
Endoscopy Centre and Endoscopy Research Institute, Zhongshan Hospital, Fudan University, Shanghai, China.
Shuo Wang
Shuo Wang
Fudan University
AI for Multi-Modal MedicineMedical Image AnalysisBiomechanics