Statistical Testing for Multiple Instance Learning via Selective Inference with Applications to Computational Pathology

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of statistical significance validation for high-attention instances in attention-based multiple instance learning, where data-dependent selection invalidates conventional hypothesis testing. To overcome this, we propose an explicit statistical framework grounded in selective inference that reformulates the evaluation of high-attention instances as a hypothesis testing problem. By introducing an adaptive reference selection mechanism to correct selection bias and compute valid p-values, our approach circumvents the limitations of traditional over-conditioning methods and achieves rigorous Type I error control. Experiments on both synthetic datasets and whole-slide images (WSIs) demonstrate that the proposed framework attains substantially higher statistical power than standard approaches, thereby providing reliable statistical guarantees for model interpretability.
📝 Abstract
Multiple instance learning (MIL) is widely used in computational pathology because it enables weakly supervised analysis of whole-slide images (WSIs) without requiring patch-level annotations. In attention-based MIL, instances with high attention scores are often interpreted as diagnostically important regions and used as visual explanations. However, attention scores alone cannot determine whether selected high-attention instances are significantly different from normal instances, limiting the reliability of attention-based explanations. In this paper, we formulate the evaluation of high-attention instances as a statistical hypothesis testing problem. Specifically, we assess whether a selected high-attention instance significantly deviates from a representative normal reference instance selected based on feature similarity. A major challenge is that both the target instance and the reference instance are selected through data-dependent procedures, rendering standard hypothesis testing invalid. To address this issue, we introduce a selective inference (SI) framework that explicitly accounts for the selection events induced by attention-based instance selection and adaptive reference selection, thereby enabling the computation of valid selective $p$-values conditional on these events. Experiments demonstrate Type-I error control on synthetic and MNIST-based data and practical applicability to pathological WSIs, with higher statistical power than the conventional over-conditioning approach.
Problem

Research questions and friction points this paper is trying to address.

Multiple Instance Learning
Selective Inference
Computational Pathology
Statistical Hypothesis Testing
Attention-based Explanation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multiple Instance Learning
Selective Inference
Computational Pathology
Statistical Hypothesis Testing
Attention Mechanism