Statistical Testing for Multiple Instance Learning via Selective Inference with Applications to Computational Pathology
This study addresses the lack of statistical significance validation for high-attention instances in attention-based multiple instance learning, where data-dependent selection invalidates conventional hypothesis testing. To overcome this, we propose an explicit statistical framework grounded in selective inference that reformulates the evaluation of high-attention instances as a hypothesis testing problem. By introducing an adaptive reference selection mechanism to correct selection bias and compute valid p-values, our approach circumvents the limitations of traditional over-conditioning methods and achieves rigorous Type I error control. Experiments on both synthetic datasets and whole-slide images (WSIs) demonstrate that the proposed framework attains substantially higher statistical power than standard approaches, thereby providing reliable statistical guarantees for model interpretability.