EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of systematic evaluation benchmarks for ophthalmic vision-language models, which hinders comprehensive assessment from disease recognition to spatial localization. We construct a unified benchmark comprising 20,000 question-answer pairs integrated from 21 datasets with deterministic annotations, covering six ocular diseases and seven question types. To transcend single-image limitations, we introduce multi-image reasoning questions constituting 44.5% of the benchmark, while generating localization gold standards by combining segmentation masks with anatomical landmarks. Under a zero-shot evaluation protocol assessing fourteen multimodal large language models, the top-performing model achieves only 62.8 points, revealing substantial deficiencies in spatial localization and cross-task generalization. This work establishes a critical diagnostic benchmark for advancing ophthalmic artificial intelligence.
📝 Abstract
Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically designed for individual diseases or specialized tasks, making it difficult to systematically evaluate whether VLMs can move beyond disease recognition toward comparative reasoning and fine-grained spatial grounding. We introduce EyeVQA, a unified visual question answering benchmark for comprehensive evaluation of ophthalmic VLMs. EyeVQA is constructed from 21 available ophthalmic datasets and contains 20,000 question-answer pairs spanning six disease groups and seven question types: Single-Choice, Multi-Select, Variable-Select, True-False, Ranking, Point Location, and Bounding Box. Gold answers are deterministically derived from source-provided diagnoses, severity grades, clinical findings, segmentation masks, bounding boxes, and anatomical landmarks, enabling reproducible evaluation without relying on model-generated annotations. Notably, 44.5% of the questions require reasoning across multiple images, extending evaluation beyond conventional single-image medical VQA. We benchmark fourteen representative general-purpose, scientific, and medically specialized VLMs under a unified zero-shot protocol. The best-performing model only achieves an overall score of 62.8, while substantial gaps remain in spatial grounding and cross-task generalization. These results highlight the limitations of current VLMs in comprehensive ophthalmic visual understanding and establish EyeVQA as a diagnostic benchmark for developing more reliable and spatially grounded ophthalmic multimodal models. The project page is available at https://github.com/PKUTHM/EyeVQA.
Problem

Research questions and friction points this paper is trying to address.

Ophthalmic Vision-Language Models
Visual Question Answering
Spatial Grounding
Benchmark Evaluation
Cross-image Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Ophthalmic VQA Benchmark
Spatial Grounding
Multi-image Reasoning
Deterministic Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Gujie Shao
Z
Zixun Xie
X
Xuechun Xing
R
Ruixiang Wang
Z
Ziyun Lan
Yanlin Qi
Yanlin Qi
University of California, Davis
Traffic PredictionsUrban ComputingData MiningGeoAI
G
Gangyi Zhang
Y
Yuxin Yang
D
Dawei Li
H
Haiming Tang