Masked Feature Encoding for Large-Scale Whole Slide Image Representation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in whole slide image analysis where intra-slide variance arising from staining heterogeneity, scanner discrepancies, and local texture differences frequently obscures critical discriminative signals. To mitigate this, we propose MFE-MIL, a multiple instance learning framework that eliminates the need for coordinate graphs or segmentation preprocessing by leveraging only raster-scan order as a weak spatial prior. The method jointly trains feature-space masking with lightweight MLP adapters to suppress redundant variance, while incorporating window-based masked reconstruction as an auxiliary regularizer. Notably, the decoder is discarded during inference to ensure computational efficiency. Extensive experiments demonstrate that MFE-MIL significantly improves classification performance across multiple datasets in terms of accuracy, F1 score, and AUC, outperforming existing coordinate-based spatial approaches. Furthermore, it substantially enhances the concordance index in survival prediction tasks.
📝 Abstract
Whole slide image (WSI) analysis in computational pathology follows a multiple instance learning (MIL) pipeline where patch embeddings are extracted independently and aggregated for slide-level prediction, but within-slide variance from staining, scanner, and local texture can overwhelm the discriminative signal. We propose Masked Feature Encoding for Multiple Instance Learning (MFE-MIL), a feature-space masking framework that trains a lightweight MLP adapter jointly with a window-based masked reconstruction branch and a MIL classification head. The two objectives are complementary. Classification guides the adapter to suppress within-slide patch variance, while window-based masked reconstruction provides an auxiliary regularizer for the adapted features without using patch coordinates, coordinate graphs, or segmentation preprocessing. The raster patch-extraction order is used only as a weak implicit prior. At inference, the decoder is removed, leaving only the adapter and MIL head. Across CAMELYON16/17, PANDA, and TCGA-BRCA with four diverse encoders, MFE-MIL improves ACC/F1 for nearly all tested aggregator-encoder settings and AUC in most, outperforms coordinate-based spatial methods (CAMIL), and achieves higher AUC than 2DMamba on three of four datasets (UNI). On five TCGA survival cohorts it improves the average concordance index for every aggregator tested, its most consistent gain. Code is available at https://github.com/AtlasAnalyticsLab/MFE-MIL.
Problem

Research questions and friction points this paper is trying to address.

Whole Slide Image
Computational Pathology
Multiple Instance Learning
Within-slide Variance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Masked Feature Encoding
Multiple Instance Learning
Computational Pathology
Whole Slide Image
Feature-space Masking
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haoyu He
Department of Computer Science and Software Engineering (CSSE), Concordia University, Montreal, Canada
B
Basile Tessier-Cloutier
McGill University, Montreal, Canada
Y
Yang Wang
Department of Computer Science and Software Engineering (CSSE), Concordia University, Montreal, Canada; Mila – Quebec AI Institute, Montreal, Canada
Mahdi S. Hosseini
Mahdi S. Hosseini
Assistant Professor, Concordia University, Mila Quebec AI Institute, McGill University
Computer VisionDeep LearningComputational Pathology