Do Pathology Vision-Language Models Truly See Pathology?

📅 2026-07-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current evaluations of pathological vision-language models (VLMs) often overlook the necessity of visual evidence and the alignment between visual and semantic representations, leading to deceptively high answer accuracy that masks fundamental deficiencies in visual understanding. To address this, this work proposes PathBind—the first multidimensional, expert-validated benchmark specifically designed to assess visual grounding capabilities. PathBind encompasses visual question answering, didactic diagram question answering, and region localization tasks, with data quality ensured through a combination of automated filtering and expert review. Experiments across 18 state-of-the-art VLMs reveal that, despite achieving high answer accuracy, these models generally exhibit weak visual-semantic binding, underscoring the urgent need to shift evaluation paradigms from merely “answering correctly” to genuinely “understanding visually.”
📝 Abstract
Pathology vision-language models (VLMs) have recently progressed rapidly and are commonly evaluated by answer accuracy on pathology VQA benchmarks. However, we dig into current evaluations and identify three overlooked issues: 1) Visual evidence is not always necessary. For instance, Gemini-3-Pro achieves 53.5% average accuracy across 5 VQA benchmarks without any visual input. 2) Domain training can improve accuracy without proportional gains in visual binding. Compared with Qwen2.5-VL-7B, Patho-R1-7B exhibits a 5.8-point lower multimodal gain and a 3.7-point lower attention IoU. 3) Entity-level attention is diffuse and weakly query-specific. On PathVG, attention maps remain highly correlated across different entity queries. These issues can lead to substantial misjudgments of pathology VLMs' actual multimodal capabilities. To this end, we present PathBind, a benchmark comprising 2,600 samples: PathBind-VQA with 1,500 questions across six dimensions, PathBind-PTA with 600 questions from a private pathology teaching atlas, and PathBind-Grounding with 500 expert-curated region-level samples. Each component undergoes task-specific automated filtering and expert review to reduce textual shortcuts and improve entity-region correspondence. We evaluate 18 representative VLMs on VQA samples of PathBind and five existing pathology VQA benchmarks, and further evaluate 10 VLMs on PathBind-Grounding and PathVG. Results show that current pathology VLMs still exhibit a substantial gap between answer-side performance and visual-semantic binding.
Problem

Research questions and friction points this paper is trying to address.

pathology vision-language models
visual-semantic binding
VQA benchmarks
multimodal evaluation
visual grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pathology Vision-Language Models
Visual-Semantic Binding
Benchmark Evaluation
Attention Analysis
Multimodal Reasoning
Chengyang Zhang
Chengyang Zhang
Sichuan University
Wenchuan Zhang
Wenchuan Zhang
Sichuan University
Clinical PathologyComputational PathologyBioinformaticsStatistics
B
Bo Li
Department of Computer Science, School of Computing, National University of Singapore
X
Xinyu Liu
College of Computer Science, Sichuan University
Jiaming Yang
Jiaming Yang
University of Michigan
Randomized Linear AlgebraOptimizationStatistics
Mengran Li
Mengran Li
Sun Yat-sen University
network scienceheterogeneous graphhypergraph
C
Chenxun Deng
Institute of Automation, Chinese Academy of Sciences
J
Jie Chen
Department of Pathology and Institute of Clinical Pathology, West China Hospital, Sichuan University
Yang Zhang
Yang Zhang
National University of Singapore
RecommendationLLM PersonalizationTrustworthy
W
Wei Ju
College of Computer Science, Sichuan University
Yuhao Yi
Yuhao Yi
Sichuan University
Optimization and ControlNetworksMachine LearningBioinformatics
H
Hong Bu
Department of Pathology and Institute of Clinical Pathology, West China Hospital, Sichuan University
Jiancheng Lv
Jiancheng Lv
University of Science and Technology of China
Operations ManagementMarketing