PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing pathological multimodal benchmarks primarily focus on final diagnostic outcomes, making it difficult to evaluate models’ fine-grained understanding of multiscale visual content in histopathology images. To address this limitation, this work proposes PathVU—the first programmable evaluation framework tailored for multiscale, fine-grained comprehension of pathological images. PathVU integrates 23 public datasets comprising 61,673 whole-slide images and over 7.25 million spatial annotations, enabling automated assessment across 14 visual question answering tasks—including localization, identification, counting, and spatial reasoning—through both region-level and whole-slide perspectives. Experiments on 18 state-of-the-art multimodal large models reveal substantial performance gaps in these fine-grained tasks, establishing PathVU as a reproducible and granular benchmark for advancing pathological multimodal understanding.
📝 Abstract
Multimodal large language models (MLLMs) are increasingly used to analyze pathology images. However, dominant multimodal benchmarks in pathology mainly score final diagnostic answers, captions, or reports. These evaluations provide limited insight into whether a model understands the multiscale visual content needed for pathology reasoning and decision-making. We introduce PathVU, a vision-anchored benchmark for fine-grained and multiscale visual understanding in computational pathology. Built from 23 public pathology imaging datasets with human-supervised labels and spatial annotations, PathVU evaluates MLLM understanding in two fields of view: Region FOV for high-resolution local regions and Slide FOV for macro whole-slide views. By converting raw annotations into deterministic task targets, PathVU enables programmatic scoring of region localization, visual recognition, quantity estimation, spatial reasoning, and insufficient-context judgment. The benchmark contains 14 VQA-style tasks, 61,673 images, and 308,070 samples across 28 organs and 7,253,526 annotations. Evaluating 18 representative general-purpose, medical-domain, and pathology-oriented MLLMs, we observe substantial limitations even in advanced models on fine-grained visual tasks across multiscale pathology images. PathVU provides a reproducible basis for developing and evaluating pathology MLLMs with explicit multiscale visual understanding.
Problem

Research questions and friction points this paper is trying to address.

multimodal large language models
pathology images
multiscale understanding
fine-grained visual understanding
computational pathology
Innovation

Methods, ideas, or system contributions that make the work stand out.

multiscale understanding
fine-grained visual reasoning
pathology image benchmark
vision-anchored evaluation
multimodal large language models