🤖 AI Summary
This study addresses the limited adaptability of frozen encoders and the low fine-tuning efficiency of pathology foundation models in whole slide image (WSI) classification by proposing a role-guided Mixture-of-Experts (MoE) framework. Methodologically, pathological role prototypes are designed to anchor expert functions, achieving expert initialization and target-domain adaptation through two-stage training. An asymmetric prototype-guided strategy is further introduced to effectively separate hard negative samples, thereby enhancing representational capacity. The proposed framework integrates knowledge distillation, Transformers, and multiple instance learning, ensuring compatibility with diverse backbones such as DINOv2 and Virchow2. Experimental results demonstrate that this approach consistently and significantly outperforms existing state-of-the-art baselines across five backbone architectures on the BRACS and PAROTID datasets.
📝 Abstract
Whole slide image classification is a fundamental task in computational pathology, where patch representation quality directly affects downstream aggregation and slide-level discriminability. Pathology foundation models are widely adopted as frozen feature extractors for WSI classification; however, their fixed encoders may produce representations insufficiently adapted to target-specific tissue patterns and discriminative cues. Fine-tuning can improve target adaptation, but introduces a trade-off between pathology-specific representation capacity and adaptation efficiency, particularly in data-scarce settings. To address this, we propose a pathology role-guided mixture-of-experts feed-forward network (MoE-FFN) framework for efficient encoder-level representation learning. We design a two-stage training paradigm to establish and adapt pathology-aware expert specialization. In source-domain expert initialization, pathology-specific priors are distilled from a frozen Virchow2 teacher into a lightweight DINOv2-small student, while role prototypes serve as weak pathological anchors to encourage distinct expert functions. MoE-FFN blocks are introduced into selected high-level transformer layers to provide transformation diversity for heterogeneous pathological patterns. In target-domain adaptation, the initialized experts are refined through asymmetric prototype-guided optimization, enhancing task-relevant positive evidence and separating confusable hard negatives. The resulting encoder extracts offline patch representations that can be directly integrated with standard MIL aggregators. Experiments on the public BRACS dataset and a private PAROTID WSI dataset across five representative backbones demonstrate consistent improvements over the strongest baseline.