See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of insufficient fine-grained perception and observational hallucinations in pathology vision-language models by proposing the ASPECT framework. Methodologically, it introduces an intermediate visual token supervision mechanism that applies explicit constraints on cellular appearance and quantity through three-stage supervised fine-tuning, feature reconstruction alignment, and reinforcement learning to enhance visual reasoning. Additionally, this work constructs PathoVernier, an expert-level benchmark dataset. Experimental results demonstrate that the proposed framework outperforms Gemini by 19.2% in accuracy on this benchmark while significantly reducing counting errors, thereby effectively advancing the fine-grained understanding capabilities for pathological images.
📝 Abstract
Pathological assessment relies on recognizing fine-grained visual details in histological images. Vision-language models (VLMs) increasingly support pathology interpretation, yet their ability to perceive these details remains inadequate. This weakness leads to inaccurate cellular observations that can persist even when final answers are correct. In this paper, we propose ASPECT to improve visually grounded reasoning through explicit supervision of cellular appearance and abundance. ASPECT trains intermediate visual tokens through pathology feature reconstruction, cell feature alignment, and count supervision. Three-stage supervised fine-tuning teaches the model to perceive, generate visual tokens, and reason, followed by reinforcement learning that rewards answer correctness and consistency with reported measurements. We also introduce PathoVernier, a benchmark of 759 expert-reviewed questions from five pathology datasets covering four cellular composition tasks. It evaluates both final answers and intermediate measurements to expose errors hidden by answer accuracy. On PathoVernier, ASPECT achieves relative accuracy gains of approximately 19.2% over the strongest baseline, Gemini-3.1-Pro, and 99.3% over its Qwen3-VL-8B backbone, while reducing RAWR, which measures counting errors within correct responses, by 28.1% and 42.7%, respectively. ASPECT also improves over its backbone on three external pathology benchmarks covering classification and question answering beyond cellular composition tasks.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Pathology
Fine-grained Perception
Visually Grounded Reasoning
Cellular Observation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visually Grounded Reasoning
Vision-Language Models
Pathology
Reinforcement Learning
Benchmark
💼 Related Jobs
No related jobs found.
Chengyang Zhang
Chengyang Zhang
Sichuan University
Wenchuan Zhang
Wenchuan Zhang
Sichuan University
Clinical PathologyComputational PathologyBioinformaticsStatistics
B
Bo Li
Department of Computer Science, School of Computing, National University of Singapore
Mengran Li
Mengran Li
Sun Yat-sen University
network scienceheterogeneous graphhypergraph
X
Xinyu Liu
College of Computer Science, Sichuan University
Jiaming Yang
Jiaming Yang
University of Michigan
Randomized Linear AlgebraOptimizationStatistics
J
Jie Chen
Department of Pathology and Institute of Clinical Pathology, West China Hospital, Sichuan University
Z
Zhang Zhang
Department of Pathology and Institute of Clinical Pathology, West China Hospital, Sichuan University
Yuhao Yi
Yuhao Yi
Sichuan University
Optimization and ControlNetworksMachine LearningBioinformatics
H
Hong Bu
Department of Pathology and Institute of Clinical Pathology, West China Hospital, Sichuan University
Jiancheng Lv
Jiancheng Lv
University of Science and Technology of China
Operations ManagementMarketing