🤖 AI Summary
This study addresses the limitation of conventional feature extraction in whole slide image classification, where patches are compressed into single embeddings that obscure fine-grained diagnostic evidence. To overcome this, we propose DI-MIL, a framework that achieves instance decoupling by clustering dense spatial tokens, thereby enhancing the attention mechanism's capacity to capture sparse evidence. Furthermore, this work introduces the first training-free post-processing module, which improves instance construction accuracy with zero additional computational overhead and establishes a new orthogonal design dimension for multiple instance learning (MIL). Extensive experiments demonstrate that DI-MIL improves 67 out of 72 metrics across four datasets, achieving an average gain of 3.64 points and outperforming seven representative baseline methods.
📝 Abstract
In weakly supervised Whole Slide Image (WSI) classification, feature extractors typically compress each image tile into a single global embedding. Consequently, slide-level aggregators are restricted to this coarse tile scale, concealing fine-grained sub-tile evidence from the attention mechanism. We introduce DI-MIL, a framework that decouples encoding context from instance granularity through decomposed instances. By clustering dense spatial tokens from a frozen foundation model within each tile, DI-MIL converts a single tile into multiple independently weighted instance embeddings. As a training-free post-encoding module, DI-MIL integrates seamlessly into existing pipelines without requiring re-encoding or downstream architectural modifications. We evaluate DI-MIL on cytopathology, a challenging testbed where sparse diagnostic signals are easily diluted within standard tiles. Across four datasets, three frozen foundation models, and two attention-based aggregators, DI-MIL demonstrates consistent efficacy, improving 67 of 72 metric-level comparisons, with the largest mean gains reaching 3.64 points under cytopathology-specific backbones. In a broader comparison against seven representative MIL baselines, DI-MIL paired with ACMIL achieves highest mean performance in 33 of 36 backbone-dataset-metric comparisons. Ablations show that direct smaller tiling inflates the extracted tile count by up to 43.3$\times$ with non-monotonic performance, whereas DI-MIL incurs zero additional image-extraction overhead while achieving the strongest overall results. These results establish instance construction as an orthogonal design dimension in MIL, supporting DI-MIL as a cost-efficient solution under sparse diagnostic evidence.