spatially-aware multiple instance learning

Designs and implements multiple-instance learning (MIL) models and aggregation mechanisms that explicitly encode spatial and distance relationships among instances, such as inter-instance distances, neighborhood connectivity, or geometry-preserving encodings. Builds representations, pooling/attention modules, and analytic pipelines that preserve and exploit spatial context across sets of patches or instances so final predictions reflect local and global spatial structure.

spatially-awaremultipleinstancelearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Conventional multiple-instance learning (MIL) methods in medical image analysis model instances—e.g., patches or slices—independently, neglecting spatial or sequential contextual dependencies, thereby limiting generalization. Method: We construct a class of synthetic classification tasks with analytically tractable optimal solutions that explicitly require models to leverage features from neighboring instances for discrimination. This enables systematic diagnosis of fundamental bottlenecks in context modeling and generalization of existing MIL approaches, including state-of-the-art relational MIL models. Contribution/Results: Through quantitative comparison against the closed-form Bayesian optimal estimator, we provide the first rigorous quantification of the substantial performance gap between mainstream MIL methods and the theoretical optimum. Experiments demonstrate that even under large-scale training, current methods fail to approach the optimal solution, underscoring the necessity of explicit context-aware mechanisms in MIL frameworks.

Correlated MIL methods struggle with optimal generalizationMIL ignores contextual relationships between instancesSynthetic task reveals generalization gaps in MIL

This work addresses the challenge of overfitting and poor generalization in multiple instance learning (MIL) under label-scarce conditions by proposing a context-based, fine-tuning-free approach. The method leverages a Perceiver architecture pretrained on diverse synthetic bag-structured datasets, integrating complementary inductive biases from varied generation strategies. This enables the model to perform accurate classification on new MIL tasks through a single forward pass with only a few labeled bags, without requiring any gradient-based adaptation. Evaluated across twelve established MIL benchmarks, the proposed approach consistently outperforms supervised baselines that rely on task-specific training, demonstrating substantially improved generalization and practical utility in few-shot MIL scenarios.

bag-structured datalow-label regimemodel adaptability

torchmil: A PyTorch-based library for deep Multiple Instance Learning

Sep 09, 2025
FM
Francisco M Castro-Macías
🏛️ University of Granada

The deep multiple instance learning (MIL) community has long suffered from a lack of standardized tooling, hindering reproducibility, fair benchmarking, and real-world deployment. To address this, we introduce *torchmil*, the first open-source PyTorch framework specifically designed for deep MIL. Guided by principles of unification, modularity, and extensibility, *torchmil* provides standardized data interfaces, curated benchmark datasets, plug-and-play model components, and a comprehensive evaluation protocol. Its key innovation is an end-to-end weakly supervised experimental pipeline that significantly lowers implementation barriers while ensuring cross-study comparability. The framework is publicly released under an open-source license and has been widely adopted in both academic research and industrial applications. Empirical validation demonstrates its effectiveness in enabling systematic, reproducible evaluation and rapid prototyping of MIL methods.

Addresses lack of reproducibility in MIL researchProvides unified framework for MIL evaluation and comparisonStandardizes deep Multiple Instance Learning model development

A Vector Symbolic Approach to Multiple Instance Learning

Nov 20, 2025
EA
Ehsan Ahmed Dhrubo
🏛️ University of Maryland, Baltimore County | North South University | CrowdStrike | Datalytica

Multi-instance learning (MIL) imposes a strict logical constraint: a bag is labeled positive *if and only if* it contains at least one positive instance. However, mainstream deep learning approaches violate this constraint, leading to inflated evaluation metrics and degraded generalization. To address this, we propose the first differentiable Vector Symbolic Architecture (VSA) framework explicitly embedding MIL’s formal logic: instances are mapped to high-dimensional symbolic vectors, and VSA algebraic operations—particularly binding and unbinding—are leveraged to explicitly encode existential quantification (“there exists”). We further introduce a learnable VSA-MaxNetwork classifier enabling end-to-end differentiable inference. Our approach uniquely unifies differentiable symbolic reasoning with deep learning, intrinsically enforcing the MIL assumption at the architectural level—thereby enhancing both interpretability and generalization. Extensive experiments on standard MIL benchmarks and medical imaging datasets demonstrate state-of-the-art performance while strictly adhering to the formal MIL definition.

Bridging raw data with symbolic representations using learned encodersEnforcing logical iff constraint in Multiple Instance Learning classificationProviding interpretable MIL framework with strict constraint adherence

Multiple Instance Verification

Jul 09, 2024
XX
Xin Xu
🏛️ University of Waikato

This work addresses the challenge of fine-grained relational modeling between a query instance and heterogeneous, relationally ambiguous target instance bags in multi-instance verification. We propose Cross-Attention Pooling (CAP), the first framework to generate query-aware bag representations. CAP introduces two novel query-guided attention functions that dynamically aggregate discriminative instances and explicitly model inter-instance dependencies—overcoming fundamental limitations of conventional MIL and Siamese architectures in capturing complex relevance structures. Evaluated on three distinct verification tasks, CAP consistently outperforms state-of-the-art MIL variants and strong baselines, achieving simultaneous gains in classification accuracy and explanation quality. Ablation studies confirm CAP’s robust capability in identifying critical instances. Overall, CAP establishes a new paradigm for interpretable multi-instance verification by unifying representation learning, relational reasoning, and explainability within a single differentiable framework.

Addressing failures of standard MIL and Siamese network adaptationsDistinguishing highly similar instances within target bagsVerifying a query instance against a heterogeneous bag of target instances

Latest Papers

What's happening recently
View more

Standard image classifiers employing global average pooling (GAP) discard spatial information, making it difficult to localize class-discriminative evidence in multi-object scenes. This work reveals, for the first time, that the conventional architecture—comprising GAP followed by a linear classification head—inherently exhibits multi-instance learning (MIL) characteristics, naturally treating an image as a bag of spatial instances. Leveraging this insight, we propose a post-hoc method that requires no model modification and recovers localized class evidence obscured by pooling through predictive grid decomposition. Experiments demonstrate that our approach effectively reconstructs faithful foreground responses using off-the-shelf classifiers and further uncovers that classification failures often stem from the intrinsic limitations of mean aggregation inherent in GAP.

global average poolingimage classificationmean aggregation

This study addresses the challenge of efficiently and accurately classifying 3D neuroimaging data (CT/MRI) under resource-constrained conditions. The authors systematically evaluate the performance of multiple-instance learning (MIL), 3D CNNs, and 3D Vision Transformers across several large-scale neuroimaging datasets and propose a novel MIL framework based on frozen, pre-trained 2D image encoders. Their findings demonstrate that a simple mean-pooling MIL approach achieves state-of-the-art performance on four out of six medium-scale tasks and remains competitive even on datasets comprising tens of thousands of scans, while training up to 25 times faster than more complex models. These results highlight the substantial efficiency and accuracy advantages of mean-pooling MIL without learnable attention mechanisms, offering a promising direction for lightweight medical image analysis.

3D Neuroimage ClassificationComputational EfficiencyModel Comparison

This work addresses a key performance bottleneck in multiple instance learning (MIL)—the reliance on linear transformations from generic to task-specific features. To overcome this limitation, the authors propose the MAMMOTH module, which introduces, for the first time, a phenotype-aware low-rank mixture-of-experts mechanism that dynamically generates lightweight, task-specific transformations for each image patch. By integrating low-rank matrix decomposition with dynamic gating, MAMMOTH enables fine-grained feature modulation with negligible parameter overhead and can be seamlessly plugged into any existing MIL architecture. Extensive evaluation across 8 MIL methods and 19 classification tasks (152 configurations in total) demonstrates consistent improvements: 130 configurations achieve higher performance, with an average accuracy gain of 3.8%. Notably, even simple aggregation strategies like mean pooling, when augmented with MAMMOTH, surpass current state-of-the-art approaches.

Computational PathologyFeature TransformationLinear Layer Bottleneck

This work addresses the challenge of high computational cost in end-to-end fine-tuning of foundation models for mammography analysis, which arises from the high resolution of mammographic images, scarce annotations, and predominantly breast-level labels. To overcome this, the authors propose MIL-PF, a framework that freezes a pretrained vision encoder, precomputes patch-level features, and introduces a lightweight attention-based multiple instance learning (MIL) aggregation module containing only 40k parameters. This design enables joint modeling of global tissue context and sparse local lesion signals without retraining large models. Evaluated on clinically scaled datasets, MIL-PF achieves state-of-the-art performance in breast cancer classification while substantially reducing training resource requirements. The code is publicly released to ensure reproducibility.

high-resolution medical imaginglimited annotationsmammography classification

This study systematically evaluates the representational capabilities of vision-language models (VLMs) and video generation models (VGMs) on spatial intelligence tasks. Using a frozen-feature probing approach, the analysis compares their performance across three dimensions: semantic labeling, instance grouping, and 3D geometric prediction. The work reveals, for the first time, a complementary relationship between VLMs and VGMs in spatial understanding: VLMs excel at semantic and instance-level recognition, whereas VGMs demonstrate superior modeling of geometric structure and camera motion dynamics. Notably, a simple fusion of their representations yields substantial gains in overall performance, simultaneously enhancing both semantic accuracy and geometric fidelity.

pretraining paradigmspatial intelligenceVideo Generation Models

Hot Scholars

UM

Urbashi Mitra

Gordon S. Marshall Chair in Engineering, University of Southern California
CommunicationsInformation TheorySignal ProcessingUnderwater Acoustic Communications
PY

Po Yang

Professor of Pervasive Intelligence at the Sheffield University
Pervasive healthcareMobile ComputingHealth Data AnalyticSmart Agriculture
JD

Jiuyang Dong

清华大学,哈尔滨工业大学
图像处理,计算病理学,显微成像
JL

Jiahan Li

PhD @ New York University, BS @ Peking University
Generative ModelingGeometric LearningAI for Science.