🤖 AI Summary
This study addresses the limited robustness of pathology foundation models to variations across centers, scanners, and staining protocols. Built upon a ViT-G/14 architecture, the proposed approach integrates DINO and iBOT self-supervised learning with a novel high-resolution Gram anchoring technique. Notably, this work is the first to establish resistance to acquisition shift as a core evaluation dimension, employing a morphology-balanced corpus for training optimization. The resulting model ranks first across 39 whole-slide-level tasks, demonstrating superior robustness against multi-source acquisition variations while achieving classification and segmentation performance on par with state-of-the-art methods. These advances substantially enhance the clinical generalizability of pathology foundation models.
📝 Abstract
Foundation models trained on large pathology image corpora now provide strong, transferable representations for computational pathology. Over the past few years a series of such models has been released, each trained on more slides than the last; on standard classification and segmentation benchmarks, the leading models are now separated by small margins. In clinical use, however, the foundation model is applied to images from hospitals, scanners, and staining protocols outside its training data. Encoders generally embed these acquisition factors alongside biological information, which may introduce downstream errors and hinder safe clinical adoption. A pathology foundation model should therefore be robust to acquisition shift without giving up representation quality, yet robustness is seldom the axis along which models are compared. In this report, we introduce HERO (Histology Encoder for Robust Representation in Oncology), a ViT-G/14 pathology foundation model trained with the DINO and iBOT objectives and refined with high-resolution Gram anchoring on a morphology-balanced corpus of 500 million tiles from approximately 575,000 clinical whole-slide images. Across the evaluated public benchmarks, HERO shows the strongest robustness to center, scanner, and stain variation among the compared state-of-the-art foundation models, performs comparably on tile-level classification, segmentation, and gene-expression prediction, ranks first on average across 39 evaluated slide-level clinical tasks, and, under an equal-weighted framework-level analysis, has the best average rank across the six benchmark frameworks.