distributional multiple instance learning

Designs, builds, or analyzes multiple-instance learning models that represent each labeled bag as a probability distribution over its instance feature embeddings instead of a single pooled vector. This includes estimating parametric or nonparametric bag-level distributions (including zero-inflated or beta-family parameterizations), aggregating patch embeddings into distributional summaries, using those distributions for bag-level prediction and uncertainty estimation, and handling sparse or zero-inflated instance signals.

distributionalmultipleinstancelearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of overfitting and poor generalization in multiple instance learning (MIL) under label-scarce conditions by proposing a context-based, fine-tuning-free approach. The method leverages a Perceiver architecture pretrained on diverse synthetic bag-structured datasets, integrating complementary inductive biases from varied generation strategies. This enables the model to perform accurate classification on new MIL tasks through a single forward pass with only a few labeled bags, without requiring any gradient-based adaptation. Evaluated across twelve established MIL benchmarks, the proposed approach consistently outperforms supervised baselines that rely on task-specific training, demonstrating substantially improved generalization and practical utility in few-shot MIL scenarios.

bag-structured datalow-label regimemodel adaptability

Which distribution were you sampled from? Towards a more tangible conception of data

Jul 24, 2024
BH
Benedikt Holtgen
🏛️ University of Tübingen | Tübingen AI Center

This paper challenges the reliance of machine learning research in the social sciences on abstract data-generating distributions, arguing that such assumptions lack empirical grounding in finite-population settings and engender interpretability and reproducibility issues. Method: The authors advocate replacing distributional assumptions with finite-population modeling, systematically advancing five core arguments grounded in statistical foundations, philosophical epistemology, and ML empirical analysis. They reconstruct the premises of learning theory by explicitly identifying the implicit assumptions and boundary conditions underlying distributional modeling. Contribution/Results: The proposed framework enhances theoretical coherence, modeling transparency, causal traceability, and practical applicability. It provides a novel paradigm and methodological foundation for sampling design, bias correction in evaluation, and reproducibility research—thereby addressing critical limitations of conventional distribution-based approaches in social-science ML applications.

Avoid assuming data-generating probability distributions in social MLChallenge reliance on abstract distributions for fairness in algorithmsPropose alternative frameworks focusing on populations not distributions

This work addresses the limitations of traditional generalization analyses, which rely on the often unverifiable assumption of independent and identically distributed (i.i.d.) data and thus struggle to accurately characterize model performance on unseen data. The paper proposes a deterministic generalization analysis framework that dispenses with any prior probabilistic assumptions. By examining the sensitivity of optimization solutions to data perturbations, it decomposes the generalization error into geometric and probabilistic components, achieving their first-ever decoupling. The framework expresses generalization bounds via a variational principle, leveraging deterministic perturbation analysis and optimization sensitivity theory to capture the discrepancy between in-sample and out-of-sample performance. Error terms are evaluated through posterior statistical hypotheses, enabling the recovery of conventional high-probability or expected generalization guarantees—all without requiring distributional assumptions.

generalizationi.i.d.optimization

This study addresses the high cost and scarcity of expert annotation for tumor proportion score (TPS) assessment in non-small cell lung cancer (NSCLC), as well as the limitations of existing multiple instance learning (MIL) approaches in handling a large number of non-expressive (zero-class) image patches. To overcome these challenges, the authors propose an end-to-end deep MIL framework that integrates CNN-based feature embedding with a multi-class network to model patch-level characteristics. Notably, they introduce, for the first time, a zero-inflated Beta distribution to probabilistically model TPS at the whole-slide level, effectively capturing both the excess of zero-class instances and the continuous nature of TPS values. Compared to linear and ridge regression baselines, the proposed method significantly improves prediction accuracy and provides calibrated confidence estimates through the concentration parameter of the learned distribution, yielding more reliable and interpretable TPS assessments.

histopathological image analysismultiple instance learningnon-small cell lung cancer

Learning Counterfactual Distributions via Kernel Nearest Neighbors

Oct 17, 2024
KC
Kyuseong Choi
🏛️ Cornell Tech | Columbia University

This paper addresses counterfactual distribution estimation and distribution matrix completion in multi-unit (e.g., users, regions)–multi-outcome (e.g., expenditure, engagement) settings, where data suffer from nonignorable missingness (MNAR), unobserved confounding, and severe sparsity—often with only a few observations per unit–outcome pair. We propose the first distributed matrix completion framework for distributional data: it leverages kernel mean embeddings to define distributional neighborhoods and enables consistent distribution recovery even under MNAR mechanisms without positivity. The method is robust to heteroscedastic noise, requiring only ≥2 observations per unit–outcome. We establish theoretical consistency of distribution recovery and enforce distributional similarity via the Maximum Mean Discrepancy (MMD). Experiments demonstrate that our approach significantly outperforms existing single-sample nearest-neighbor and standard matrix completion methods under sparse, biased-sampling, and heteroscedastic regimes.

Estimating distributions from finite samples using kernel nearest neighborsHandling unobserved confounding and heteroscedastic noise robustlyLearning multivariate distributions with missing not at random data

Latest Papers

What's happening recently
View more

Prior Distribution and Model Confidence

Sep 05, 2025
MK
Maksim Kazanskii
🏛️ Independent Researcher | BIOPOLIS/CIBIO

This work investigates how training data distribution affects the generalization of image classification models and proposes a model-agnostic, retraining-free confidence assessment framework. The method fuses multiple embeddings into a unified embedding space to jointly characterize the structure of the training distribution; it then employs an adaptive distance metric to quantify each sample’s deviation from this distribution, enabling confidence calibration, low-confidence prediction filtering, and out-of-distribution (OOD) detection. The framework is architecture-agnostic and domain-transferable, delivering consistent improvements across diverse backbones—including ResNet and Vision Transformers—without architectural modification or fine-tuning. Notably, accuracy gains are especially pronounced after filtering low-confidence predictions. Empirically, it enhances classification robustness and reliability under distribution shifts, offering a lightweight, plug-and-play, distribution-aware inference mechanism for trustworthy AI systems.

Filtering low-confidence predictions using embedding distanceFramework for model confidence without retrainingImpact of training data distribution on model performance

A Vector Symbolic Approach to Multiple Instance Learning

Nov 20, 2025
EA
Ehsan Ahmed Dhrubo
🏛️ University of Maryland, Baltimore County | North South University | CrowdStrike | Datalytica

Multi-instance learning (MIL) imposes a strict logical constraint: a bag is labeled positive *if and only if* it contains at least one positive instance. However, mainstream deep learning approaches violate this constraint, leading to inflated evaluation metrics and degraded generalization. To address this, we propose the first differentiable Vector Symbolic Architecture (VSA) framework explicitly embedding MIL’s formal logic: instances are mapped to high-dimensional symbolic vectors, and VSA algebraic operations—particularly binding and unbinding—are leveraged to explicitly encode existential quantification (“there exists”). We further introduce a learnable VSA-MaxNetwork classifier enabling end-to-end differentiable inference. Our approach uniquely unifies differentiable symbolic reasoning with deep learning, intrinsically enforcing the MIL assumption at the architectural level—thereby enhancing both interpretability and generalization. Extensive experiments on standard MIL benchmarks and medical imaging datasets demonstrate state-of-the-art performance while strictly adhering to the formal MIL definition.

Bridging raw data with symbolic representations using learned encodersEnforcing logical iff constraint in Multiple Instance Learning classificationProviding interpretable MIL framework with strict constraint adherence

Standard Bagging ensembles often suffer from overconfidence and redundancy due to uniform voting weights that ignore the varying local competencies of base learners. This work proposes the SCSB framework, which unifies ensemble pruning and probability calibration into a joint optimization problem over the probability simplex. By minimizing out-of-bag loss with an added concave quadratic sparsity-inducing penalty, SCSB overcomes the theoretical limitation of the L1 norm—which fails to induce sparsity on the simplex—while preserving model-agnosticism. The method achieves compression rates up to 96%, substantially reduces expected calibration error, yields linear inference speedup, and maintains or even improves generalization accuracy in most cases.

baggingensemble learningmodel calibration

The randomness in the number of distinct observations within bootstrap resamples—termed sample diversity—contributes to the variance of out-of-bag (OOB) error estimation, yet its isolated impact remains poorly quantified. Method: To disentangle this effect, we introduce Sequential Bootstrap into the OOB analysis framework, enabling precise, explicit control over the number of unique observations per resample—thereby decoupling diversity variability from other sources of stochasticity. We evaluate it under Breiman’s classic five-OBB experimental design on both synthetic and real-world datasets, using multiple random seeds. Contribution/Results: Our approach significantly reduces OOB estimation variance without compromising model predictive accuracy, empirically confirming that sample diversity fluctuations constitute a primary source of OOB variance. This work provides the first reproducible, controllable empirical pathway for systematic variance decomposition in bootstrap-based ensembles, thereby advancing both the theoretical understanding and practical toolkit for uncertainty quantification in ensemble learning.

Evaluating Sequential Bootstrap's impact on variance reduction in ensemblesInvestigating how distinct observation variability affects OOB error estimationProviding empirical evidence for variance decomposition in bootstrap systems

Conventional multiple-instance learning (MIL) methods in medical image analysis model instances—e.g., patches or slices—independently, neglecting spatial or sequential contextual dependencies, thereby limiting generalization. Method: We construct a class of synthetic classification tasks with analytically tractable optimal solutions that explicitly require models to leverage features from neighboring instances for discrimination. This enables systematic diagnosis of fundamental bottlenecks in context modeling and generalization of existing MIL approaches, including state-of-the-art relational MIL models. Contribution/Results: Through quantitative comparison against the closed-form Bayesian optimal estimator, we provide the first rigorous quantification of the substantial performance gap between mainstream MIL methods and the theoretical optimum. Experiments demonstrate that even under large-scale training, current methods fail to approach the optimal solution, underscoring the necessity of explicit context-aware mechanisms in MIL frameworks.

Correlated MIL methods struggle with optimal generalizationMIL ignores contextual relationships between instancesSynthetic task reveals generalization gaps in MIL

Hot Scholars

RC

Rama Chellappa

Bloomberg Distinguished Professor, Johns Hopkins University
Image Analysisartificial intelligencebiometricsComputer Vision
LC

Lei Cao

Assistant Professor, University of Arizona/Research Scientist, MIT CSAIL
DatabasesMachine learning
EM

Effrosyni Mavroudi

Research Scientist, FAIR, Meta AI
Computer VisionMachine LearningVideo Understanding
PP

Prateek Prasanna

Associate Professor, Stony Brook University
Medical VisionBiomedical image analysisRadiogenomicsRadiomics
ML

Min-Ling Zhang

Professor, School of Computer Science and Engineering, Southeast University, China
Artificial IntelligenceMachine LearningData Mining