One Tile, Multiple Instances: Rethinking MIL for Sparse Diagnostic Evidence

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of conventional feature extraction in whole slide image classification, where patches are compressed into single embeddings that obscure fine-grained diagnostic evidence. To overcome this, we propose DI-MIL, a framework that achieves instance decoupling by clustering dense spatial tokens, thereby enhancing the attention mechanism's capacity to capture sparse evidence. Furthermore, this work introduces the first training-free post-processing module, which improves instance construction accuracy with zero additional computational overhead and establishes a new orthogonal design dimension for multiple instance learning (MIL). Extensive experiments demonstrate that DI-MIL improves 67 out of 72 metrics across four datasets, achieving an average gain of 3.64 points and outperforming seven representative baseline methods.
📝 Abstract
In weakly supervised Whole Slide Image (WSI) classification, feature extractors typically compress each image tile into a single global embedding. Consequently, slide-level aggregators are restricted to this coarse tile scale, concealing fine-grained sub-tile evidence from the attention mechanism. We introduce DI-MIL, a framework that decouples encoding context from instance granularity through decomposed instances. By clustering dense spatial tokens from a frozen foundation model within each tile, DI-MIL converts a single tile into multiple independently weighted instance embeddings. As a training-free post-encoding module, DI-MIL integrates seamlessly into existing pipelines without requiring re-encoding or downstream architectural modifications. We evaluate DI-MIL on cytopathology, a challenging testbed where sparse diagnostic signals are easily diluted within standard tiles. Across four datasets, three frozen foundation models, and two attention-based aggregators, DI-MIL demonstrates consistent efficacy, improving 67 of 72 metric-level comparisons, with the largest mean gains reaching 3.64 points under cytopathology-specific backbones. In a broader comparison against seven representative MIL baselines, DI-MIL paired with ACMIL achieves highest mean performance in 33 of 36 backbone-dataset-metric comparisons. Ablations show that direct smaller tiling inflates the extracted tile count by up to 43.3$\times$ with non-monotonic performance, whereas DI-MIL incurs zero additional image-extraction overhead while achieving the strongest overall results. These results establish instance construction as an orthogonal design dimension in MIL, supporting DI-MIL as a cost-efficient solution under sparse diagnostic evidence.
Problem

Research questions and friction points this paper is trying to address.

Whole Slide Image classification
Multiple Instance Learning
sparse diagnostic evidence
weakly supervised learning
cytopathology
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multiple Instance Learning
Whole Slide Image
Decomposed Instances
Token Clustering
Training-free
R
Runsheng Liu
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong, China.
Cheng Jin
Cheng Jin
Ph.D. Student, School of Computer Science and Engineering, HKUST
Knowledge DistillationComputational PathologyAI for Science
Hao Jiang
Hao Jiang
The Hong Kong University of Science and Technology
Medical Image AnalysisDeep LearningComputational Pathology
H
Hao Chen
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong, China.