Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of pathology foundation models to acquisition-related confounders in cross-institutional tissue similarity assessment, where reliance on shortcut features compromises robustness. To this end, we propose MOSAIC, a benchmark framework integrating relative similarity evaluation with multi-institutional datasets to systematically compare general-purpose multimodal large language models (MLLMs) against specialized pathology models. Our results demonstrate that MLLMs significantly outperform conventional encoders through joint semantic-visual reasoning, and that mere data scaling fails to mitigate domain shift limitations inherent in the latter. This work validates the robustness of MLLMs in cross-institutional retrieval, establishes their potential as viable alternatives to pathology foundation models, and introduces a novel paradigm for enhancing quality control capabilities in multicenter computational pathology.
📝 Abstract
State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons cross slide or institution boundaries. We show that general-purpose multimodal LLMs, without being trained as pathology foundation models, consistently outperform these specialized models in cross-domain histological similarity judgments. Using a relative similarity framework that we release as the MOSAIC (Model Similarity Assessment across Institutions and Cohorts) benchmark, we evaluate 17 models across 6 datasets and find that pathology encoders often rank same-institution, different-disease tiles as more similar than same-disease, different-institution tiles, a clinically dangerous failure mode invisible to standard within-domain evaluations. LLMs appear less susceptible to this failure, likely because they perform semantic visual comparison of morphology and tissue architecture rather than relying on shortcut features tied to acquisition context. Scaling training data does not resolve the problem for pathology encoders, implicating the learning objective rather than data coverage. Our results expose a fundamental robustness gap in current pathology foundation models and establish multimodal LLMs as a viable alternative for cross-institutional retrieval, dataset harmonization, and multi-site quality control. Code and data will be released upon acceptance.
Problem

Research questions and friction points this paper is trying to address.

pathology foundation models
cross-domain histological similarity
multimodal LLMs
robustness gap
shortcut features
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal LLMs
Pathology Foundation Models
Cross-Domain Similarity
MOSAIC Benchmark
Shortcut Features
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yishu Zhang
University of North Carolina at Chapel Hill
Yun Li
Yun Li
University of North Carolina
Statistical GeneticsBioinformaticsGenomicsComplex Traits
D
Daiwei Zhang
University of North Carolina at Chapel Hill