PathLang: A Language-Centered Benchmark for Vision-Language Models in Computational Pathology
This study addresses the insufficient robustness of pathology vision-language models to variations in diagnostic phrasing. For the first time, language is treated as an independent variable axis to construct a language-centric zero-shot benchmark. By fixing images and labels while systematically perturbing terminology, specificity, and reporting style, this work evaluates models across classification, retrieval, and open-vocabulary diagnosis tasks, with prompts validated by six board-certified pathologists. The findings reveal that model performance is highly sensitive to clinically equivalent paraphrasing and that image-text alignment quality does not necessarily translate into inter-class separability. To facilitate future research, the complete evaluation resources are publicly released.