Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the vulnerability of foundation models in computational pathology to non-biological confounders—such as staining and scanning variations—during cross-center deployment, which leads representations to rely on spurious shortcut features and compromises generalization. The authors propose CRoMa (Cross-confounder Robustness Margin), a sample-level distributional robustness metric that quantifies robustness as a marginal distribution property across the entire cohort by comparing representation distances between biologically matched samples and those perturbed by the same confounding factor, eschewing conventional count-based aggregation. CRoMa reveals, for the first time, intrinsic heterogeneity in model robustness and establishes a Pareto trade-off framework between typical performance and tail robustness. Extensive validation across multiple slide- and patch-level histopathology benchmarks demonstrates that CRoMa consistently ranks models, uncovers a confounder-dominated robustness lower tail in all encoders, and shows that higher CRoMa values effectively predict reduced shortcut-induced performance degradation after fine-tuning.
📝 Abstract
Pathology foundation models are approaching clinical deployment, yet remain vulnerable to systematic non-biological variation across centres. Differences in tissue preparation, staining and scanning are strongly encoded in their representations, enabling shortcut learning and weakening generalisation across cohorts and institutions. The Robustness Index (RI) quantifies whether local representation geometry is dominated by biology or by non-biological variation, but its count-based formulation discards distance information. We show that adding distance weights changes little because the deeper limitation lies in RI's pooled, fixed-neighbourhood design, which obscures sample-level heterogeneity and effectively evaluates only a model-dependent subset of samples. We introduce the Cross-confounder Robustness Margin (CRoMa), a sample-resolved measure that directly compares distances to cross-confounder biological matches and same-confounder biological distractors. CRoMa recasts robustness as a cohort-wide margin distribution rather than a single pooled score. We evaluated frozen representations from 20 tile-level encoders across three benchmarks and 4 slide-level encoders on a fourth. Rankings by median CRoMa were broadly consistent across datasets, while the underlying distributions revealed substantial within-model heterogeneity. Every tile encoder retained a confounder-dominated lower tail, whose prevalence and severity varied markedly across models. These distinct robustness profiles frame model selection as a Pareto trade-off between typical and lower-tail robustness. Higher CRoMa was also associated with smaller shortcut-induced performance drops after supervised adaptation. By turning representation geometry into a distributional robustness readout that anticipates downstream shortcut susceptibility, CRoMa provides a principled basis for robustness assessment and model selection.
Problem

Research questions and friction points this paper is trying to address.

distributional robustness
pathology foundation models
non-biological variation
shortcut learning
confounder
Innovation

Methods, ideas, or system contributions that make the work stand out.

distributional robustness
pathology foundation models
shortcut learning
representation geometry
confounder-aware evaluation
🔎 Similar Papers
No similar papers found.