imbalance-aware active learning

Designs and evaluates active learning strategies that explicitly incorporate prior information about class frequencies or model beliefs to prioritize sampling of under-represented classes and select informative, class-balanced examples so minority-class performance improves while annotation cost is reduced. This work builds acquisition functions or sampling policies that trade off informativeness and class-frequency priors, estimate or integrate priors (e.g., from a pre-trained model), and analyze the impact on labeling budget and class-wise performance.

imbalance-awareactivelearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.27
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Real-world image-text data often suffer from class imbalance and label noise, which severely degrade model performance on minority classes. This work presents the first systematic study of active learning under this dual challenge and introduces a collaborative active learning framework that integrates priors from foundation models. By leveraging imbalance-aware collaborative decision-making between a foundation model and a lightweight model, the approach enables efficient and robust sample querying. Extensive experiments on multiple cross-modal imbalanced datasets demonstrate that the proposed method significantly enhances robustness to label noise while reducing annotation costs by over 50% without compromising model performance.

active learningannotation efficiencyclass imbalance

To address the challenge of harmonizing diversity and uncertainty sampling in active learning, this paper proposes TCM—a two-stage dynamic sampling framework tailored for self-supervised pretraining. In Stage I, TypiClust-based clustering ensures initial sample diversity to mitigate cold-start issues; in Stage II, margin-based uncertainty sampling is activated to enhance discriminative capability. Self-supervised pretraining strengthens representation robustness, while a learnable stage-switching mechanism enables data-volume-adaptive scheduling. Evaluated across multiple benchmark datasets, TCM consistently outperforms state-of-the-art methods under both low-resource (<10% labeled data) and high-resource (>50% labeled data) settings. It achieves superior sampling quality and generalization stability, offering a scalable, pretraining-aware paradigm for modern active learning.

Diversity in SelectionMachine LearningUncertainty Handling

Transductive Active Learning: Theory and Applications

Feb 13, 2024
JH
Jonas Hubotter
🏛️ ETH Zurich

This work addresses active learning in realistic settings where sampling is constrained to an accessible region, while prediction targets may lie outside this domain. To tackle this challenge, we propose a transductive active learning framework tailored to real-world prediction objectives, leveraging adaptive uncertainty minimization for efficient out-of-domain inference. Theoretically, we establish, for the first time under general regularity assumptions, the uniform consistency of the decision rule—guaranteeing convergence to the minimal uncertainty achievable by accessible data—providing strong and broadly applicable statistical guarantees. Methodologically, our approach integrates transductive learning, Bayesian optimization, and large language model (LLM) fine-tuning. Empirical evaluations demonstrate substantial improvements in sample efficiency on LLM active fine-tuning and safety-critical Bayesian optimization tasks, achieving state-of-the-art performance.

Applications in neural networks and Bayesian optimizationGeneralization of active learningMinimizing uncertainty in predictions

DIRECT: Deep Active Learning under Imbalance and Label Noise

Dec 14, 2023
SN
Shyam Nuggehalli
🏛️ University of Wisconsin | University of Washington

To address sample selection bias arising from the coexistence of class imbalance and label noise in deep active learning, this paper proposes a robust one-dimensional threshold-driven active learning paradigm. The method jointly models class imbalance and label noise—first achieved in active learning—and employs deep feature embedding followed by one-dimensional projection to robustly estimate an inter-class separation threshold. This threshold defines a priority region near the decision boundary where high-uncertainty samples are selected for labeling. The framework is theoretically compatible with batch querying and label-noise tolerance. Evaluated on multiple imbalanced benchmark datasets, it reduces annotation cost by over 60% compared to state-of-the-art active learning methods and improves accuracy by more than 80% relative to random sampling, while significantly enhancing minority-class recognition performance.

Addresses class imbalance impact on minority class performanceHandles label noise and reduces annotation costs significantlyProposes active learning to collect balanced informative examples

Re-Benchmarking Pool-Based Active Learning for Binary Classification

Jun 15, 2023
PL
Po-Yi Lu
🏛️ National Taiwan University | Google Cloud AI Research

This paper addresses the longstanding controversy surrounding the effectiveness of Uncertainty Sampling (US) in active learning. Methodologically, it establishes a transparent, reproducible, and modular open-source pool-based active learning benchmark framework. It systematically re-evaluates mainstream query strategies on binary classification tasks, corrects configuration biases present in prior benchmarks, and—crucially—first identifies and quantifies “model compatibility”: the substantial degradation in US performance arising from inconsistency between the querying and training models. Key contributions include: (i) empirically demonstrating that US remains competitive across most benchmark datasets; (ii) rectifying several previously misattributed conclusions about strategy efficacy; and (iii) providing a standardized PyTorch-based experimental platform that integrates multiple query strategies, base models, and datasets, supporting automated evaluation and rigorous statistical analysis. This benchmark has since become the de facto standard in the active learning community.

Evaluates Uncertainty Sampling's competitiveness in Active Learning for tabular dataIdentifies model compatibility as a key factor affecting Uncertainty Sampling's performanceProvides a comprehensive benchmark to compare Active Learning strategies effectively

Latest Papers

What's happening recently
View more

Optimal Labeler Assignment and Sampling for Active Learning in the Presence of Imperfect Labels

Dec 14, 2025
PA
Pouya Ahadi
🏛️ Georgia Institute of Technology | Ford Motor Company

To address high label noise in active learning caused by annotator ability disparities—particularly erroneous labeling of complex instances—this paper proposes a robust, noise-resilient active learning framework. Methodologically: (1) it formulates an optimal annotator allocation model grounded in game theory, minimizing the worst-case potential noise per iteration; (2) it introduces an uncertainty-aware, noise-robust sampling strategy; and (3) it integrates multi-annotator confidence-weighted ensemble learning with noise-robust loss modeling. Extensive experiments across multiple benchmark datasets demonstrate an average 5.2% improvement in classification accuracy and a 37% reduction in label-noise sensitivity, significantly outperforming state-of-the-art active learning methods. The core contribution lies in the first unified integration of annotator capability modeling, noise-aware sampling, and robust ensemble learning within a closed-loop active learning pipeline.

Minimize label noise in active learning cyclesOptimally assign query points to labelersReduce impact of imperfect labels on classifier

This study investigates why active learning struggles to outperform random sampling in extremely low-resource neural machine translation settings with only 100–500 training samples. It systematically evaluates the core assumption underlying active learning—that informativeness and diversity effectively guide sample selection—under such data-scarce conditions. Through comprehensive experiments employing multiple active learning strategies, random baselines, and rigorously controlled variables, the work reveals for the first time that the optimization objectives of active learning exhibit no significant correlation with final test performance. Instead, the order of training samples and their interaction with pretraining data emerge as more critical factors influencing model effectiveness. These findings challenge prevailing active learning paradigms and suggest new directions for designing sampling strategies in few-shot machine translation scenarios.

active learningfew-shot learningmachine translation

This work proposes a novel active learning paradigm that eliminates the need for an initially pretrained candidate model, thereby circumventing the substantial computational overhead and procedural complexity of conventional approaches. Instead, the method directly employs randomly initialized CNN and Transformer models for sample selection, guided by high-confidence (HC), low-confidence (LC), and hybrid HCLC sampling strategies. Experimental results demonstrate that the LC strategy consistently achieves superior performance across most scenarios, and the proposed framework attains accuracy comparable to traditional methods reliant on pretrained models—yet with markedly enhanced efficiency, flexibility, and practicality across multiple benchmark datasets.

active learningcandidate modelsdeep learning

This work addresses the challenge in active learning where reliance on randomly initialized seed sets limits the ability to leverage related datasets for reducing annotation costs. To overcome this, the authors propose Active-Transfer Bagging (ATBagging), a novel approach that uniquely integrates transfer learning with Bagging ensembles. ATBagging estimates information gain via Bayesian predictive distributions and incorporates diversity in feature space through determinantal point processes (DPPs) for sample selection. By jointly optimizing both informativeness and diversity during both seed set construction and subsequent query stages, the method achieves consistent and significant improvements over existing techniques across four real-world datasets—QM9, ERA5, Forbes 2000, and Beijing PM2.5—particularly under low labeling budgets, where it substantially increases the area under the learning curve.

active learningdata acquisitionlabeling cost

This study addresses the cognitive bias in evaluating active feature acquisition strategies caused by uneven coverage in offline data, which conflates uncertainty arising from missing data (epistemic) with uninformative features (aleatoric). We propose an evaluation framework based on Prior-data Fitted Networks (PFNs), whose core innovation lies in replacing total predictive entropy with posterior expected aleatoric entropy as the reward function. This effectively decouples epistemic and aleatoric uncertainties, thereby eliminating evaluation bias. By integrating tabular foundation models with active feature learning techniques, the proposed method significantly reduces value estimation bias on both synthetic and real-world datasets. Furthermore, it generates confidence intervals with high empirical coverage and substantially enhances downstream policy selection performance.

Active Feature AcquisitionEpistemic BiasOffline Data

Hot Scholars

XQ

Xihe Qiu

Associate Professor, Shanghai University of Engineering Science
AI for HealthcareVision-Language ModelsReinforcement LearningLarge Language Models
LL

Liang Liu

Highly Cited Researcher, EEE Department, The Hong Kong Polytechnic University
convex optimizationMIMOInternet of Thingsintegrated sensing and communication
FL

Fan Lyu

NLPR, CASIA
Computer VisionMachine LearningArtificial Intelligence
BG

Bin Guo

Professor of Computer Science, Northwestern Polytechnical University
Ubiquitous ComputingMobile Crowd Sensing
KM

Kleanthis Malialis

KIOS Research and Innovation Center of Excellence, University of Cyprus
Machine LearningData Stream MiningIncremental LearningConcept Drift