clip-wise multiple instance learning

Design and implement models that partition long sequential data into short clips, compute per-clip features or scores, and aggregate those clip-level outputs into a sequence-level prediction using multiple instance learning (MIL) pooling or attention mechanisms. Build clip-wise feature extractors, MIL aggregation modules, and training/evaluation procedures that handle weak labels, variable numbers of clips, and robust sequence-level predictions such as counts or residuals.

clip-wisemultipleinstancelearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.52
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a key performance bottleneck in multiple instance learning (MIL)—the reliance on linear transformations from generic to task-specific features. To overcome this limitation, the authors propose the MAMMOTH module, which introduces, for the first time, a phenotype-aware low-rank mixture-of-experts mechanism that dynamically generates lightweight, task-specific transformations for each image patch. By integrating low-rank matrix decomposition with dynamic gating, MAMMOTH enables fine-grained feature modulation with negligible parameter overhead and can be seamlessly plugged into any existing MIL architecture. Extensive evaluation across 8 MIL methods and 19 classification tasks (152 configurations in total) demonstrates consistent improvements: 130 configurations achieve higher performance, with an average accuracy gain of 3.8%. Notably, even simple aggregation strategies like mean pooling, when augmented with MAMMOTH, surpass current state-of-the-art approaches.

Computational PathologyFeature TransformationLinear Layer Bottleneck

torchmil: A PyTorch-based library for deep Multiple Instance Learning

Sep 09, 2025
FM
Francisco M Castro-Macías
🏛️ University of Granada

The deep multiple instance learning (MIL) community has long suffered from a lack of standardized tooling, hindering reproducibility, fair benchmarking, and real-world deployment. To address this, we introduce *torchmil*, the first open-source PyTorch framework specifically designed for deep MIL. Guided by principles of unification, modularity, and extensibility, *torchmil* provides standardized data interfaces, curated benchmark datasets, plug-and-play model components, and a comprehensive evaluation protocol. Its key innovation is an end-to-end weakly supervised experimental pipeline that significantly lowers implementation barriers while ensuring cross-study comparability. The framework is publicly released under an open-source license and has been widely adopted in both academic research and industrial applications. Empirical validation demonstrates its effectiveness in enabling systematic, reproducible evaluation and rapid prototyping of MIL methods.

Addresses lack of reproducibility in MIL researchProvides unified framework for MIL evaluation and comparisonStandardizes deep Multiple Instance Learning model development

This work addresses the challenge of overfitting and poor generalization in multiple instance learning (MIL) under label-scarce conditions by proposing a context-based, fine-tuning-free approach. The method leverages a Perceiver architecture pretrained on diverse synthetic bag-structured datasets, integrating complementary inductive biases from varied generation strategies. This enables the model to perform accurate classification on new MIL tasks through a single forward pass with only a few labeled bags, without requiring any gradient-based adaptation. Evaluated across twelve established MIL benchmarks, the proposed approach consistently outperforms supervised baselines that rely on task-specific training, demonstrating substantially improved generalization and practical utility in few-shot MIL scenarios.

bag-structured datalow-label regimemodel adaptability

Multiple Instance Verification

Jul 09, 2024
XX
Xin Xu
🏛️ University of Waikato

This work addresses the challenge of fine-grained relational modeling between a query instance and heterogeneous, relationally ambiguous target instance bags in multi-instance verification. We propose Cross-Attention Pooling (CAP), the first framework to generate query-aware bag representations. CAP introduces two novel query-guided attention functions that dynamically aggregate discriminative instances and explicitly model inter-instance dependencies—overcoming fundamental limitations of conventional MIL and Siamese architectures in capturing complex relevance structures. Evaluated on three distinct verification tasks, CAP consistently outperforms state-of-the-art MIL variants and strong baselines, achieving simultaneous gains in classification accuracy and explanation quality. Ablation studies confirm CAP’s robust capability in identifying critical instances. Overall, CAP establishes a new paradigm for interpretable multi-instance verification by unifying representation learning, relational reasoning, and explainability within a single differentiable framework.

Addressing failures of standard MIL and Siamese network adaptationsDistinguishing highly similar instances within target bagsVerifying a query instance against a heterogeneous bag of target instances

CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling

Sep 28, 2024
JZ
Jihai Zhang
🏛️ The Chinese University of Hong Kong | Shanghai AI Laboratory | Schoow University

CLIP suffers from significant fine-grained visual information loss due to its monolithic feature encoding. To address this, we propose a model-agnostic Diversified Multi-Expert Upgrade (DMU) strategy—the first to integrate sparsely activated Mixture-of-Experts (MoE) into the CLIP architecture. Leveraging parameter sharing (with independent feed-forward networks only) and dynamic routing, DMU efficiently distills a single dense CLIP checkpoint into multiple complementary expert submodels, yielding a zero-adaptation, plug-and-play CLIP-MoE. Crucially, DMU requires no modifications to downstream frameworks and supports direct, seamless upgrade of any pre-trained CLIP checkpoint. Evaluated on zero-shot image–text retrieval, image classification, and MLLM visual encoding tasks, DMU consistently improves performance by 3.2–7.8% while increasing computational overhead by less than 5%.

Develop cost-effective method for creating diverse CLIP modelsEnhance CLIP's feature diversity to reduce information lossOptimize model capacity and computational cost via CLIP-MoE

Latest Papers

What's happening recently
View more

This study investigates the pronounced performance degradation of large language models (LLMs) when processing multi-instance inputs, a phenomenon that intensifies with increasing instance count—termed “performance collapse.” Through systematic evaluation of mainstream LLMs under controlled variations of context length and instance number, the work identifies instance quantity as the primary driver of performance decline, exerting a stronger effect than context length alone. Empirical results reveal that model performance begins to mildly deteriorate at 20–100 instances and subsequently undergoes sharp deterioration at larger scales. These findings offer novel empirical insights and a foundational understanding for advancing multi-instance reasoning mechanisms and guiding future model optimization strategies.

Context LengthInstance CountLarge Language Models

This work addresses the catastrophic forgetting in CLIP-based class-incremental learning, which arises from attribute extraction and aggregation shifts due to reliance solely on current-task data. To mitigate this issue, the authors propose a two-stage decoupled framework: first, principal geodesic analysis is employed to anchor class-specific attributes in the hyperspherical embedding space, stabilizing feature extraction; second, lightweight task-specific experts regularized by a variational information bottleneck are introduced, with inference routed via an optimal transport mechanism. This approach effectively alleviates forgetting and significantly outperforms state-of-the-art methods across multiple benchmarks, thereby enhancing the incremental learning capability of the base CLIP model.

Attribute AggregationAttribute ExtractionCatastrophic Forgetting

Standard image classifiers employing global average pooling (GAP) discard spatial information, making it difficult to localize class-discriminative evidence in multi-object scenes. This work reveals, for the first time, that the conventional architecture—comprising GAP followed by a linear classification head—inherently exhibits multi-instance learning (MIL) characteristics, naturally treating an image as a bag of spatial instances. Leveraging this insight, we propose a post-hoc method that requires no model modification and recovers localized class evidence obscured by pooling through predictive grid decomposition. Experiments demonstrate that our approach effectively reconstructs faithful foreground responses using off-the-shelf classifiers and further uncovers that classification failures often stem from the intrinsic limitations of mean aggregation inherent in GAP.

global average poolingimage classificationmean aggregation

This study addresses the challenge of efficiently and accurately classifying 3D neuroimaging data (CT/MRI) under resource-constrained conditions. The authors systematically evaluate the performance of multiple-instance learning (MIL), 3D CNNs, and 3D Vision Transformers across several large-scale neuroimaging datasets and propose a novel MIL framework based on frozen, pre-trained 2D image encoders. Their findings demonstrate that a simple mean-pooling MIL approach achieves state-of-the-art performance on four out of six medium-scale tasks and remains competitive even on datasets comprising tens of thousands of scans, while training up to 25 times faster than more complex models. These results highlight the substantial efficiency and accuracy advantages of mean-pooling MIL without learnable attention mechanisms, offering a promising direction for lightweight medical image analysis.

3D Neuroimage ClassificationComputational EfficiencyModel Comparison

Hot Scholars

YH

Yuyang Huang

University of Chicago
system for mlcomputer systemoperating system
TS

Tom Sander

Meta FAIR & Ecole polytechnique
Privacy Preserving Machine Learning
AD

Alain Durmus

Ecole polytechnique
Machine learningStatistics
CS

Chengchao Shen

Central South University
Computer VisionMachine Learning