Score
Design and implement models that partition long sequential data into short clips, compute per-clip features or scores, and aggregate those clip-level outputs into a sequence-level prediction using multiple instance learning (MIL) pooling or attention mechanisms. Build clip-wise feature extractors, MIL aggregation modules, and training/evaluation procedures that handle weak labels, variable numbers of clips, and robust sequence-level predictions such as counts or residuals.
This work addresses a key performance bottleneck in multiple instance learning (MIL)—the reliance on linear transformations from generic to task-specific features. To overcome this limitation, the authors propose the MAMMOTH module, which introduces, for the first time, a phenotype-aware low-rank mixture-of-experts mechanism that dynamically generates lightweight, task-specific transformations for each image patch. By integrating low-rank matrix decomposition with dynamic gating, MAMMOTH enables fine-grained feature modulation with negligible parameter overhead and can be seamlessly plugged into any existing MIL architecture. Extensive evaluation across 8 MIL methods and 19 classification tasks (152 configurations in total) demonstrates consistent improvements: 130 configurations achieve higher performance, with an average accuracy gain of 3.8%. Notably, even simple aggregation strategies like mean pooling, when augmented with MAMMOTH, surpass current state-of-the-art approaches.
The deep multiple instance learning (MIL) community has long suffered from a lack of standardized tooling, hindering reproducibility, fair benchmarking, and real-world deployment. To address this, we introduce *torchmil*, the first open-source PyTorch framework specifically designed for deep MIL. Guided by principles of unification, modularity, and extensibility, *torchmil* provides standardized data interfaces, curated benchmark datasets, plug-and-play model components, and a comprehensive evaluation protocol. Its key innovation is an end-to-end weakly supervised experimental pipeline that significantly lowers implementation barriers while ensuring cross-study comparability. The framework is publicly released under an open-source license and has been widely adopted in both academic research and industrial applications. Empirical validation demonstrates its effectiveness in enabling systematic, reproducible evaluation and rapid prototyping of MIL methods.
This work addresses the challenge of overfitting and poor generalization in multiple instance learning (MIL) under label-scarce conditions by proposing a context-based, fine-tuning-free approach. The method leverages a Perceiver architecture pretrained on diverse synthetic bag-structured datasets, integrating complementary inductive biases from varied generation strategies. This enables the model to perform accurate classification on new MIL tasks through a single forward pass with only a few labeled bags, without requiring any gradient-based adaptation. Evaluated across twelve established MIL benchmarks, the proposed approach consistently outperforms supervised baselines that rely on task-specific training, demonstrating substantially improved generalization and practical utility in few-shot MIL scenarios.
This work addresses the challenge of fine-grained relational modeling between a query instance and heterogeneous, relationally ambiguous target instance bags in multi-instance verification. We propose Cross-Attention Pooling (CAP), the first framework to generate query-aware bag representations. CAP introduces two novel query-guided attention functions that dynamically aggregate discriminative instances and explicitly model inter-instance dependencies—overcoming fundamental limitations of conventional MIL and Siamese architectures in capturing complex relevance structures. Evaluated on three distinct verification tasks, CAP consistently outperforms state-of-the-art MIL variants and strong baselines, achieving simultaneous gains in classification accuracy and explanation quality. Ablation studies confirm CAP’s robust capability in identifying critical instances. Overall, CAP establishes a new paradigm for interpretable multi-instance verification by unifying representation learning, relational reasoning, and explainability within a single differentiable framework.
CLIP suffers from significant fine-grained visual information loss due to its monolithic feature encoding. To address this, we propose a model-agnostic Diversified Multi-Expert Upgrade (DMU) strategy—the first to integrate sparsely activated Mixture-of-Experts (MoE) into the CLIP architecture. Leveraging parameter sharing (with independent feed-forward networks only) and dynamic routing, DMU efficiently distills a single dense CLIP checkpoint into multiple complementary expert submodels, yielding a zero-adaptation, plug-and-play CLIP-MoE. Crucially, DMU requires no modifications to downstream frameworks and supports direct, seamless upgrade of any pre-trained CLIP checkpoint. Evaluated on zero-shot image–text retrieval, image classification, and MLLM visual encoding tasks, DMU consistently improves performance by 3.2–7.8% while increasing computational overhead by less than 5%.
This study investigates the pronounced performance degradation of large language models (LLMs) when processing multi-instance inputs, a phenomenon that intensifies with increasing instance count—termed “performance collapse.” Through systematic evaluation of mainstream LLMs under controlled variations of context length and instance number, the work identifies instance quantity as the primary driver of performance decline, exerting a stronger effect than context length alone. Empirical results reveal that model performance begins to mildly deteriorate at 20–100 instances and subsequently undergoes sharp deterioration at larger scales. These findings offer novel empirical insights and a foundational understanding for advancing multi-instance reasoning mechanisms and guiding future model optimization strategies.
This work addresses the catastrophic forgetting in CLIP-based class-incremental learning, which arises from attribute extraction and aggregation shifts due to reliance solely on current-task data. To mitigate this issue, the authors propose a two-stage decoupled framework: first, principal geodesic analysis is employed to anchor class-specific attributes in the hyperspherical embedding space, stabilizing feature extraction; second, lightweight task-specific experts regularized by a variational information bottleneck are introduced, with inference routed via an optimal transport mechanism. This approach effectively alleviates forgetting and significantly outperforms state-of-the-art methods across multiple benchmarks, thereby enhancing the incremental learning capability of the base CLIP model.
Standard image classifiers employing global average pooling (GAP) discard spatial information, making it difficult to localize class-discriminative evidence in multi-object scenes. This work reveals, for the first time, that the conventional architecture—comprising GAP followed by a linear classification head—inherently exhibits multi-instance learning (MIL) characteristics, naturally treating an image as a bag of spatial instances. Leveraging this insight, we propose a post-hoc method that requires no model modification and recovers localized class evidence obscured by pooling through predictive grid decomposition. Experiments demonstrate that our approach effectively reconstructs faithful foreground responses using off-the-shelf classifiers and further uncovers that classification failures often stem from the intrinsic limitations of mean aggregation inherent in GAP.
This study addresses the challenge of efficiently and accurately classifying 3D neuroimaging data (CT/MRI) under resource-constrained conditions. The authors systematically evaluate the performance of multiple-instance learning (MIL), 3D CNNs, and 3D Vision Transformers across several large-scale neuroimaging datasets and propose a novel MIL framework based on frozen, pre-trained 2D image encoders. Their findings demonstrate that a simple mean-pooling MIL approach achieves state-of-the-art performance on four out of six medium-scale tasks and remains competitive even on datasets comprising tens of thousands of scans, while training up to 25 times faster than more complex models. These results highlight the substantial efficiency and accuracy advantages of mean-pooling MIL without learnable attention mechanisms, offering a promising direction for lightweight medical image analysis.