PathLang: A Language-Centered Benchmark for Vision-Language Models in Computational Pathology

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient robustness of pathology vision-language models to variations in diagnostic phrasing. For the first time, language is treated as an independent variable axis to construct a language-centric zero-shot benchmark. By fixing images and labels while systematically perturbing terminology, specificity, and reporting style, this work evaluates models across classification, retrieval, and open-vocabulary diagnosis tasks, with prompts validated by six board-certified pathologists. The findings reveal that model performance is highly sensitive to clinically equivalent paraphrasing and that image-text alignment quality does not necessarily translate into inter-class separability. To facilitate future research, the complete evaluation resources are publicly released.
📝 Abstract
Pathology vision-language models (VLMs) have shown strong visual perception ability, but their robustness in the language domain remains poorly characterized. Existing pathology VLM benchmarks largely rely on canonical closed-set prompts or perturb only generic templates, treating language as a fixed evaluation component rather than a variable axis of model behavior. In clinical practice, however, diagnostic language varies across reports, institutions, and candidate diagnoses. We introduce PathLang, a language-centered and clinically grounded zero-shot benchmark. PathLang holds the underlying slides, ground-truth labels, and image-text evaluation direction fixed while systematically varying only the diagnostic language, so that performance differences reflect how a diagnosis is phrased rather than what is imaged. The language variation follows how pathologists actually rephrase diagnoses (terminology, specificity, and reporting style), and all prompts and candidate pools are validated by six board-certified pathologists. PathLang covers four task families: (1) zero-shot classification with image-text alignment analysis, (2) cross-modal retrieval, (3) paraphrase robustness, including semantic-equivalence paraphrases, length and reporting-style variation, and prompt ensembling, and (4) open-vocabulary diagnosis retrieval over four candidate pools with distinct forms of semantic competition. Across nine VLMs and five public datasets spanning four organs, we find that performance is highly sensitive to clinically equivalent paraphrases, varies substantially across forms of semantic competition, and that image-text alignment quality does not necessarily translate into inter-class separability. We release the prompt corpus, candidate pools, pre-computed text embeddings, and evaluation code at https://anonymous.4open.science/r/PathLang.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Computational Pathology
Language Robustness
Benchmark
Zero-shot Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Computational Pathology
Zero-Shot Benchmark
Paraphrase Robustness
Open-Vocabulary Retrieval
💼 Related Jobs
No related jobs found.
F
Fanqi Cheng
New York University, New York, NY, USA
K
Kuo Gong
Columbia University, New York, NY, USA
S
Shangke Liu
Weill Cornell Medicine, New York, NY, USA
B
Beidi Zhao
University of British Columbia, Vancouver, BC, Canada
Junchao Zhu
Junchao Zhu
Vanderbilt University
Z
Zheyu Zhu
University of Pennsylvania, Philadelphia, PA, USA
L
Leiyue Zhao
Johns Hopkins University, Baltimore, MD, USA
Fengbei Liu
Fengbei Liu
Cornell University
Computer VisionMedical Image Analysis
J
John Cannon
New York Medical College, New York, NY, USA
G
Gang Wang
University of British Columbia, Vancouver, BC, Canada
Z
Zu-hua Gao
BC Cancer Agency, Vancouver, BC, Canada
Kenji Ikemura
Kenji Ikemura
Weill Cornell Medicine - New York Presbyterian
Biomedical EngineeringMolecular PathologyClinical Informatics
Yihe Yang
Yihe Yang
Northwell Health
Renal PathologyAnatomic and Clinical PathologyEpidemiology and BiostatisticsNephrologyClinical Research
Yaohong Wang
Yaohong Wang
The University of Texas MD Anderson Cancer Center, Houston, TX, USA
Yuankai Huo
Yuankai Huo
Computer Science, Vanderbilt University
Medical Image AnalysisDeep LearningData Mining
X
Xiaoxiao Li
University of British Columbia, Vancouver, BC, Canada
Mert R. Sabuncu
Mert R. Sabuncu
Cornell University, Cornell Tech, Weill Cornell Medicine
AI for medical imagingmedical image computingmedical image analysismachine learning
Ruining Deng
Ruining Deng
Weill Cornell Medicine
Medical Image AnalysisDeep LearningDigital Pathology