margin-based selective sampling

Designs and evaluates active learning and selective-labeling procedures that compute a prediction margin (the difference between top predicted scores) for unlabeled instances and rank or select low-margin, high-uncertainty examples to query for labels. Builds sampling policies, label-collection workflows, and analytic evaluations that minimize annotation effort under a label budget while preserving classifier performance metrics such as accuracy and AUC.

margin-basedselectivesampling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

DIRECT: Deep Active Learning under Imbalance and Label Noise

Dec 14, 2023
SN
Shyam Nuggehalli
🏛️ University of Wisconsin | University of Washington

To address sample selection bias arising from the coexistence of class imbalance and label noise in deep active learning, this paper proposes a robust one-dimensional threshold-driven active learning paradigm. The method jointly models class imbalance and label noise—first achieved in active learning—and employs deep feature embedding followed by one-dimensional projection to robustly estimate an inter-class separation threshold. This threshold defines a priority region near the decision boundary where high-uncertainty samples are selected for labeling. The framework is theoretically compatible with batch querying and label-noise tolerance. Evaluated on multiple imbalanced benchmark datasets, it reduces annotation cost by over 60% compared to state-of-the-art active learning methods and improves accuracy by more than 80% relative to random sampling, while significantly enhancing minority-class recognition performance.

Addresses class imbalance impact on minority class performanceHandles label noise and reduces annotation costs significantlyProposes active learning to collect balanced informative examples

Active Learning For Contextual Linear Optimization: A Margin-Based Approach

May 11, 2023
ML
Mo Liu
🏛️ University of California, Berkeley

This work addresses contextual linear optimization, where label acquisition—i.e., querying target function coefficients—is costly, and aims to minimize decision error measured by the Smart Predict-then-Optimize (SPO) loss via active learning. We propose the first SPO-loss-driven active learning framework, which directly incorporates the SPO loss into the query strategy and introduces a novel margin-based selection criterion grounded in distance-to-decision-boundary analysis to dynamically identify the most informative unlabeled instances. Theoretically, we establish the first label complexity upper bound and generalization risk guarantee for SPO and its surrogate SPO+ loss under margin conditions. Empirically, our method substantially reduces labeling effort while outperforming fully supervised baselines on personalized pricing and shortest-path prediction tasks.

Active LearningDecision Loss MinimizationLinear Optimization

In open-set active learning, unlabeled data often contain out-of-distribution (OOD) samples; blind annotation incurs substantial cost waste. Existing methods struggle to jointly optimize sample informativeness and in-distribution (ID) purity, and heavily rely on OOD labels or auxiliary training. This paper proposes CLIPNAL—a novel active learning framework that eliminates reliance on labeled OOD data for the first time. It leverages a pre-trained CLIP model to perform unsupervised OOD detection and filtering, then selects highly informative samples exclusively from the remaining ID pool. CLIPNAL jointly enhances purity and informativeness via semantic-vision alignment scoring and a two-stage selection strategy (purity-first followed by informativeness-based refinement), requiring neither OOD annotations nor additional model training. Experiments across diverse open-set settings demonstrate that CLIPNAL achieves state-of-the-art model performance at the lowest annotation cost, significantly outperforming existing approaches.

Balances informativeness and purity of unlabeled samplesMinimizes wasted annotation costs in open-set active learningReduces dependence on out-of-distribution (OOD) samples

Contextual Active Model Selection

Jul 13, 2022
XL
Xuefeng Liu
🏛️ University of Chicago | Argonne National Laboratory

To address the high labeling cost in scenarios where pre-trained models coexist with abundant unlabeled data, this paper proposes an online, context-aware active model selection framework. In each round, it dynamically selects the optimal pre-trained model based on input context for prediction and adaptively decides whether to query the true label. The method integrates context modeling, multi-armed bandit theory, active learning, and online optimization—introducing, for the first time, a context-driven model selection mechanism and a co-optimized active querying strategy. It establishes the first theoretical guarantees on regret and query complexity under both adversarial and stochastic environments. Empirical evaluation on benchmarks including CIFAR-10 and DRIFT demonstrates that the approach achieves comparable or superior accuracy using less than 10% of the labeled data required by state-of-the-art methods.

Minimize labeling costs for pre-trained models.Reduce label requests with strategic querying.Select best model adaptively using context.

Transductive Active Learning: Theory and Applications

Feb 13, 2024
JH
Jonas Hubotter
🏛️ ETH Zurich

This work addresses active learning in realistic settings where sampling is constrained to an accessible region, while prediction targets may lie outside this domain. To tackle this challenge, we propose a transductive active learning framework tailored to real-world prediction objectives, leveraging adaptive uncertainty minimization for efficient out-of-domain inference. Theoretically, we establish, for the first time under general regularity assumptions, the uniform consistency of the decision rule—guaranteeing convergence to the minimal uncertainty achievable by accessible data—providing strong and broadly applicable statistical guarantees. Methodologically, our approach integrates transductive learning, Bayesian optimization, and large language model (LLM) fine-tuning. Empirical evaluations demonstrate substantial improvements in sample efficiency on LLM active fine-tuning and safety-critical Bayesian optimization tasks, achieving state-of-the-art performance.

Applications in neural networks and Bayesian optimizationGeneralization of active learningMinimizing uncertainty in predictions

Latest Papers

What's happening recently
View more

This work addresses the challenge of selecting an appropriate active learning strategy in medical image classification, where suboptimal choices can significantly increase annotation costs and must be made before exhausting limited labeling budgets. To this end, the authors propose the ALDA framework, which employs a brief pilot phase—using only 15–30% of the total budget—to fit learning curves of candidate strategies and predict the annotation effort required to meet clinical performance targets. ALDA introduces a “deployment window” to quantify sensitivity to uncertainty in clinical thresholds and integrates risk-aware decision rules to recommend the optimal strategy. By reframing strategy selection as a deployment-oriented decision problem, ALDA jointly optimizes annotation cost and robustness. Experiments demonstrate that ALDA accurately identifies the best-performing strategy, reducing annotation costs by up to 82% compared to poor alternatives.

active learningannotation costclinical performance

This study addresses a critical limitation in conventional active learning, which assumes annotators are perfectly reliable—an assumption that rarely holds in real-world settings where labels are often noisy or missing. To bridge this gap, the authors present the first large-scale empirical evaluation of eight state-of-the-art deep active learning algorithms under realistic annotation conditions, using genuine crowdsourced text labels rather than synthetically injected noise. The assessment spans three standard datasets and systematically examines algorithmic robustness and practicality in the presence of both label noise and abstention (i.e., annotators declining to label). The findings reveal substantial performance disparities among existing methods when deployed in authentic environments, offering crucial guidance for real-world applications. To support further research, the authors publicly release their collected dataset of real human annotations.

active learningcrowd-sourced annotationsnoisy oracles

This study investigates why active learning struggles to outperform random sampling in extremely low-resource neural machine translation settings with only 100–500 training samples. It systematically evaluates the core assumption underlying active learning—that informativeness and diversity effectively guide sample selection—under such data-scarce conditions. Through comprehensive experiments employing multiple active learning strategies, random baselines, and rigorously controlled variables, the work reveals for the first time that the optimization objectives of active learning exhibit no significant correlation with final test performance. Instead, the order of training samples and their interaction with pretraining data emerge as more critical factors influencing model effectiveness. These findings challenge prevailing active learning paradigms and suggest new directions for designing sampling strategies in few-shot machine translation scenarios.

active learningfew-shot learningmachine translation

This study addresses the prediction performance bottleneck caused by missing auxiliary information during deployment and limited annotation budgets. To overcome this, it proposes ALCATRAs, a unified framework that integrates data acquisition, surrogate construction, and downstream prediction. Through task selection strategies and surrogate learning, the framework adaptively allocates resources under cost constraints to acquire critical auxiliary information, leveraging active learning, multi-task learning, and surrogate model transfer for efficient optimization. Theoretically, the authors demonstrate that ALCATRAs effectively reduces the prediction error bound. Empirical evaluations on benchmarks such as the UCI Heart Disease dataset confirm that the proposed approach significantly improves both sample efficiency and predictive accuracy compared to baseline methods.

Deployment AsymmetryMissing-by-DesignMulti-Task Active Learning

Hot Scholars

RF

Richard F. Lyon

Research Scientist, Google Inc.
Machine HearingSignal ProcessingImage SensorsPhotography
PG

Patrick Gallinari

Professor Sorbonne University / Criteo AI Lab
Machine LearningDeep LearningPhysics-aware Deep LearningNatural Language Processing
CU

Chamira U. S. Edussooriya

University of Moratuwa
Multi-dimensional Signal ProcessingDigital FiltersLight FieldsGraph Signal Processing
XZ

Xuekai Zhu

Shanghai Jiao Tong University
Synthetic DataReasoningLanguage Model
ZA

Zeynep Akata

Professor at Technical University of Munich and Director at Helmholtz Munich
Machine LearningVision and LanguageZero-Shot Learning