dual-stream multiple-instance learning

Design and implement dual‑stream multiple‑instance learning systems that operate on bags of local instances with parallel visual and text streams: compute text‑conditioned local instance scores, aggregate those into global bag‑ or scan‑level representations, and build training and inference mechanisms that supervise the global stream (including with external model–derived labels) while enforcing representational consistency and alignment between the two streams.

dual-streammultiple-instancelearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.44
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of overfitting and poor generalization in multiple instance learning (MIL) under label-scarce conditions by proposing a context-based, fine-tuning-free approach. The method leverages a Perceiver architecture pretrained on diverse synthetic bag-structured datasets, integrating complementary inductive biases from varied generation strategies. This enables the model to perform accurate classification on new MIL tasks through a single forward pass with only a few labeled bags, without requiring any gradient-based adaptation. Evaluated across twelve established MIL benchmarks, the proposed approach consistently outperforms supervised baselines that rely on task-specific training, demonstrating substantially improved generalization and practical utility in few-shot MIL scenarios.

bag-structured datalow-label regimemodel adaptability

Standard image classifiers employing global average pooling (GAP) discard spatial information, making it difficult to localize class-discriminative evidence in multi-object scenes. This work reveals, for the first time, that the conventional architecture—comprising GAP followed by a linear classification head—inherently exhibits multi-instance learning (MIL) characteristics, naturally treating an image as a bag of spatial instances. Leveraging this insight, we propose a post-hoc method that requires no model modification and recovers localized class evidence obscured by pooling through predictive grid decomposition. Experiments demonstrate that our approach effectively reconstructs faithful foreground responses using off-the-shelf classifiers and further uncovers that classification failures often stem from the intrinsic limitations of mean aggregation inherent in GAP.

global average poolingimage classificationmean aggregation

Multiple Instance Verification

Jul 09, 2024
XX
Xin Xu
🏛️ University of Waikato

This work addresses the challenge of fine-grained relational modeling between a query instance and heterogeneous, relationally ambiguous target instance bags in multi-instance verification. We propose Cross-Attention Pooling (CAP), the first framework to generate query-aware bag representations. CAP introduces two novel query-guided attention functions that dynamically aggregate discriminative instances and explicitly model inter-instance dependencies—overcoming fundamental limitations of conventional MIL and Siamese architectures in capturing complex relevance structures. Evaluated on three distinct verification tasks, CAP consistently outperforms state-of-the-art MIL variants and strong baselines, achieving simultaneous gains in classification accuracy and explanation quality. Ablation studies confirm CAP’s robust capability in identifying critical instances. Overall, CAP establishes a new paradigm for interpretable multi-instance verification by unifying representation learning, relational reasoning, and explainability within a single differentiable framework.

Addressing failures of standard MIL and Siamese network adaptationsDistinguishing highly similar instances within target bagsVerifying a query instance against a heterogeneous bag of target instances

A Vector Symbolic Approach to Multiple Instance Learning

Nov 20, 2025
EA
Ehsan Ahmed Dhrubo
🏛️ University of Maryland, Baltimore County | North South University | CrowdStrike | Datalytica

Multi-instance learning (MIL) imposes a strict logical constraint: a bag is labeled positive *if and only if* it contains at least one positive instance. However, mainstream deep learning approaches violate this constraint, leading to inflated evaluation metrics and degraded generalization. To address this, we propose the first differentiable Vector Symbolic Architecture (VSA) framework explicitly embedding MIL’s formal logic: instances are mapped to high-dimensional symbolic vectors, and VSA algebraic operations—particularly binding and unbinding—are leveraged to explicitly encode existential quantification (“there exists”). We further introduce a learnable VSA-MaxNetwork classifier enabling end-to-end differentiable inference. Our approach uniquely unifies differentiable symbolic reasoning with deep learning, intrinsically enforcing the MIL assumption at the architectural level—thereby enhancing both interpretability and generalization. Extensive experiments on standard MIL benchmarks and medical imaging datasets demonstrate state-of-the-art performance while strictly adhering to the formal MIL definition.

Bridging raw data with symbolic representations using learned encodersEnforcing logical iff constraint in Multiple Instance Learning classificationProviding interpretable MIL framework with strict constraint adherence

Learning from Data Streams: An Overview and Update

Dec 30, 2022
JR
Jesse Read
🏛️ École Polytechnique | IP-Paris | University of Helsinki

Existing data stream learning research often relies on unrealistic assumptions—such as single-pass processing and strict online constraints—leading to ill-defined problem formulations, biased evaluation protocols, and misalignment with industrial requirements. Method: This paper systematically critiques and deconstructs these restrictive assumptions, proposing a “de-paradigmized” framework that centers modeling on concept drift and temporal dependence while abandoning rigid formal stream constraints; algorithmic design integrates time-series analysis, concept drift detection, robust statistical learning, and privacy-preserving techniques—rejecting isolated development of bespoke streaming algorithms. Contribution/Results: The work yields a methodology guide grounded in industrial practice, fostering renewed consensus between academia and industry. It significantly enhances model robustness, interpretability, and privacy compliance in real-world dynamic environments.

Addressing unrealistic assumptions in data-stream learning algorithmsRedefining supervised data-stream learning with concept driftShifting focus to robustness, privacy, and interpretability in streams

Latest Papers

What's happening recently
View more

This work addresses critical limitations of existing Mamba-based multiple instance learning approaches in whole-slide image analysis, which often disrupt two-dimensional spatial locality, struggle to model fine-grained cellular structures, and incur excessive peak memory consumption during inference. To overcome these challenges, the authors propose MambaBack, a hybrid architecture that innovatively integrates Gated CNNs with BiMamba2. By leveraging Hilbert space-filling curve sampling, the method preserves local spatial structure and establishes a hierarchical local–global modeling mechanism. Furthermore, an asymmetric chunking strategy decouples training and inference processes to substantially reduce memory overhead. Evaluated on five benchmark datasets, MambaBack consistently outperforms seven state-of-the-art methods, achieving superior classification performance while significantly lowering deployment memory requirements.

Mambamemory efficiencyMultiple Instance Learning

This study addresses the performance degradation of self-supervised learning on continuous video streams caused by high intra-batch sample similarity, providing the first systematic analysis of streaming training bottlenecks. Methodologically, it constructs the WT++ dataset and proposes the StreamMAE framework, which optimizes the input pipeline via motion-biased cropping and introduces stream-aware regularization to improve the streaming adaptation of masked autoencoders (MAE). Experimental results demonstrate that the proposed approach surpasses existing streaming baselines, achieves performance comparable to independent and identically distributed (i.i.d.) training, and exhibits continuous improvement as data duration increases.

Continuous Video StreamsIntra-batch SimilaritySelf-Supervised Learning

Existing vision Mixture-of-Experts (MoE) approaches perform routing at the image or patch level, which struggles to align with the instance-centric nature of object detection. This work proposes a hierarchical instance-conditioned MoE architecture that introduces a two-stage routing mechanism—operating at both scene and instance levels—within a DETR-style detector, achieving fine-grained expert assignment aligned with instance queries for the first time. The method employs a lightweight scene router and an instance router that jointly balance sparse computation with the heterogeneity of individual instances. Experiments demonstrate that the model outperforms dense DINO baselines and simplified routing variants on COCO, significantly enhancing small object detection performance and offering preliminary evidence of functional specialization among experts.

expert specializationinstance-level routingMixture-of-Experts

This study addresses the challenge of efficiently and accurately classifying 3D neuroimaging data (CT/MRI) under resource-constrained conditions. The authors systematically evaluate the performance of multiple-instance learning (MIL), 3D CNNs, and 3D Vision Transformers across several large-scale neuroimaging datasets and propose a novel MIL framework based on frozen, pre-trained 2D image encoders. Their findings demonstrate that a simple mean-pooling MIL approach achieves state-of-the-art performance on four out of six medium-scale tasks and remains competitive even on datasets comprising tens of thousands of scans, while training up to 25 times faster than more complex models. These results highlight the substantial efficiency and accuracy advantages of mean-pooling MIL without learnable attention mechanisms, offering a promising direction for lightweight medical image analysis.

3D Neuroimage ClassificationComputational EfficiencyModel Comparison

This study addresses the limitations of traditional instance-based learning, which relies on handcrafted rules and struggles to generalize novel patterns under distribution shift and high-dimensional data. To overcome these challenges, this work proposes a neural instance-based learning framework that generalizes symbolic matching to neural embedding similarity, integrating structured memory mechanisms to optimize retrieval and utility fusion. Furthermore, it introduces a surprise memory to isolate weakly matched samples and designs an observation-driven graduation algorithm that promotes novel patterns into active memory, thereby enhancing drift resistance in sequential decision-making. Experimental results demonstrate that the proposed method improves accuracy by 6 to 17 percentage points across five task categories, significantly enhancing adaptability in cold-start, anomaly detection, and concept drift scenarios.

Concept driftHigh-dimensional sequential dataInstance-Based Learning

Hot Scholars

ZW

Zixing Wang

Purdue University
Embodied AIRobotics
DK

Devesh K. Jha

Senior Principal Research Scientist, MERL
Machine LearningRoboticsArtificial IntelligenceMotion Planning
DJ

Deepali Jain

Google Deepmind
Artificial IntelligenceRoboticsReinforcement Learning
AP

Anil Parwani

Professor of Pathology and Biomedical Informatics
Artificial Intelligencepathology informaticsprostate cancerkidney cancer
WC

Wei Chen

The Ohio State University
Pathology