🤖 AI Summary
This study addresses the challenge in whole slide image analysis where intra-slide variance arising from staining heterogeneity, scanner discrepancies, and local texture differences frequently obscures critical discriminative signals. To mitigate this, we propose MFE-MIL, a multiple instance learning framework that eliminates the need for coordinate graphs or segmentation preprocessing by leveraging only raster-scan order as a weak spatial prior. The method jointly trains feature-space masking with lightweight MLP adapters to suppress redundant variance, while incorporating window-based masked reconstruction as an auxiliary regularizer. Notably, the decoder is discarded during inference to ensure computational efficiency. Extensive experiments demonstrate that MFE-MIL significantly improves classification performance across multiple datasets in terms of accuracy, F1 score, and AUC, outperforming existing coordinate-based spatial approaches. Furthermore, it substantially enhances the concordance index in survival prediction tasks.
📝 Abstract
Whole slide image (WSI) analysis in computational pathology follows a multiple instance learning (MIL) pipeline where patch embeddings are extracted independently and aggregated for slide-level prediction, but within-slide variance from staining, scanner, and local texture can overwhelm the discriminative signal. We propose Masked Feature Encoding for Multiple Instance Learning (MFE-MIL), a feature-space masking framework that trains a lightweight MLP adapter jointly with a window-based masked reconstruction branch and a MIL classification head. The two objectives are complementary. Classification guides the adapter to suppress within-slide patch variance, while window-based masked reconstruction provides an auxiliary regularizer for the adapted features without using patch coordinates, coordinate graphs, or segmentation preprocessing. The raster patch-extraction order is used only as a weak implicit prior. At inference, the decoder is removed, leaving only the adapter and MIL head. Across CAMELYON16/17, PANDA, and TCGA-BRCA with four diverse encoders, MFE-MIL improves ACC/F1 for nearly all tested aggregator-encoder settings and AUC in most, outperforms coordinate-based spatial methods (CAMIL), and achieves higher AUC than 2DMamba on three of four datasets (UNI). On five TCGA survival cohorts it improves the average concordance index for every aggregator tested, its most consistent gain. Code is available at https://github.com/AtlasAnalyticsLab/MFE-MIL.