Can Multimodal Large Language Models Understand OCT?

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current evaluations of OCT image understanding are largely confined to coarse-grained classification or isolated question-answering tasks, failing to comprehensively assess the full spectrum of cognitive processes from visual perception to clinical reasoning. To address this gap, this work proposes OCT-Bench—the first fine-grained, multi-level clinical cognition benchmark for OCT image understanding—encompassing 20 tasks across three dimensions: perception, cognition, and reasoning. These tasks evaluate critical capabilities including anatomical structure recognition, lesion characterization, spatial relationships, and diagnostic decision-making. Built upon 4,137 OCT images from seven public datasets and 10,076 high-quality multiple-choice questions, OCT-Bench establishes a standardized evaluation framework. Systematic assessment of 20 representative multimodal large language models reveals significant shortcomings in reliably interpreting OCT images, with neither medical adaptation nor model scaling consistently improving performance across all cognitive levels.
📝 Abstract
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding to coarse-grained disease classification or isolated visual question answering, leaving the complete cognitive process from visual perception to clinical reasoning insufficiently evaluated. To address this limitation, we introduce OCT-Bench, a comprehensive benchmark dedicated to OCT image understanding. OCT-Bench comprises 10,076 high-quality multiple-choice questions constructed from 4,137 OCT images across seven public datasets. Following the real-world clinical interpretation workflow, we establish a hierarchical capability taxonomy consisting of 20 fine-grained tasks across three dimensions: Perception, Cognition, and Reasoning. These tasks cover a broad range of capabilities, including imaging attributes, retinal anatomy, lesion characteristics, spatial relationships, disease assessment, therapeutic decision-making, and prognostic management. We systematically evaluate 20 representative MLLMs, including proprietary models, open-source general-purpose models, and medical-domain models. Experimental results demonstrate that current models remain substantially short of reliable OCT understanding. Moreover, neither medical-domain adaptation nor increased model scale consistently improves performance across capability levels. OCT-Bench enables comprehensive and fine-grained evaluation of MLLMs, providing a foundation for identifying capability bottlenecks and advancing clinically grounded OCT understanding.
Problem

Research questions and friction points this paper is trying to address.

multimodal large language models
optical coherence tomography
medical image understanding
clinical reasoning
benchmark evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

OCT-Bench
multimodal large language models
fine-grained evaluation
clinical reasoning
hierarchical capability taxonomy
🔎 Similar Papers