balanced batch sampling

Designs and implements batch-sampling algorithms that construct training mini-batches with controlled class proportions (e.g., class-balanced or batch-balanced) by selecting samples from current data and optionally from stored or high-confidence historical examples; these procedures reduce bias toward dominant classes and stabilize model updates during training or online adaptation.

balancedbatchsampling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a self-balancing sequential sampling method tailored for applications such as auditing, scheduling, and sampling, where both unpredictability and distributional convergence are critical. By adaptively adjusting selection probabilities, the method achieves optimal $O(n^{-1})$ convergence of the empirical distribution to the target distribution while preserving maximal sample unpredictability—substantially improving upon the $O(n^{-1/2})$ rate of conventional i.i.d. sampling. Theoretical analysis reveals that this approach uniquely solves an entropy-regularized optimization problem and exhibits a diffusion limit linked to the Ornstein-Uhlenbeck process. Empirical results confirm its effectiveness in minimizing repeated selections and coverage gaps, all while rigorously maintaining control over unpredictability.

empirical distribution convergencesampling biasself-balancing

Understanding and Mitigating the Bias in Sample Selection for Learning with Noisy Labels

Jan 24, 2024
QW
Qi Wei
🏛️ Nanyang Technological University | Zhejiang University

In noisy label learning, sample selection suffers from dual biases: data bias (imbalanced selection sets) and training bias (error accumulation). To address these issues, this paper proposes ITEM, a noise-tolerant expert model. ITEM is the first to jointly model and mitigate both biases in a unified framework. It introduces a lightweight multi-expert robust network architecture, integrated with a dual-weighted class-discriminative sampler and a hybrid mini-batch training strategy. Additionally, an error-robust optimization mechanism is incorporated to enhance generalization under label noise. Extensive experiments on multiple benchmark datasets with synthetic and real-world label noise demonstrate that ITEM consistently outperforms state-of-the-art methods, achieving average accuracy gains of 3–5% while reducing parameter count by over 20%. The source code is publicly available.

Addresses bias in sample selection for noisy labelsMitigates data and training bias in selection methodsProposes ITEM model for debiased learning and robust performance

Generalization Bounds for Dependent Data using Online-to-Batch Conversion

May 22, 2024
SC
Sagnik Chatterjee
🏛️ Indraprastha Institute of Information Technology | Delhi (IIIT -D)

This work addresses the problem of deriving generalization error upper bounds for batch learning algorithms under mixing stochastic processes (i.e., dependent data), without imposing any stability assumptions on the batch learner. The method introduces a novel analytical framework based on Online-to-Batch conversion: stability requirements are shifted to an associated online learner, and a new notion of online algorithm stability—defined via the first-order Wasserstein distance—is proposed for the first time. It is shown that the Exponentially Weighted Average (EWA) algorithm satisfies this stability condition. Consequently, the framework yields both expectation- and high-probability generalization bounds applicable to *any* batch learning algorithm; under i.i.d. assumptions, the bounds reduce to classical forms up to correction terms governed by the mixing decay rate. The resulting bounds are explicitly computable, substantially broadening the applicability and practical utility of learning theory under data dependence.

Dependent data sourcesGeneralization error boundsOnline-to-Batch conversion framework

TS-RSR: A provably efficient approach for batch bayesian optimization

Mar 07, 2024
ZR
Zhaolin Ren
🏛️ Harvard University

To address the challenge of jointly maximizing predictive mean, uncertainty, and minimizing intra-batch redundancy in batch Bayesian optimization (BO), this paper proposes Thompson Sampling with Regret-to-Sigma Ratio (TS-RSR). TS-RSR is the first method to incorporate the regret-to-sigma ratio into the batch acquisition objective, leveraging a Thompson sampling approximation to explicitly balance exploration and exploitation across batch points. We provide theoretical guarantees of convergence. By integrating Gaussian process modeling with an efficient batch sampling scheme, TS-RSR significantly reduces intra-batch redundancy. Extensive experiments on synthetic benchmarks and real-world tasks demonstrate that TS-RSR consistently outperforms state-of-the-art batch BO methods, achieving new state-of-the-art performance.

Develops TS-RSR for efficient batch Bayesian OptimizationMinimizes redundancy while targeting high-value or uncertain pointsProvides theoretical guarantees and outperforms benchmark algorithms

Latest Papers

What's happening recently
View more

This study addresses the combinatorial optimization problem of identifying a balanced sampling design from a large population under a fixed inclusion probability, such that the weighted estimator of an auxiliary variable closely approximates the known population total—a task of exponential complexity. To tackle this challenge, the authors propose a heuristic approach based on genetic algorithms, which iteratively refines the sampling scheme by integrating minimum support designs with candidate samples exhibiting high balance. This method overcomes the limitations of the traditional cube method in achieving balance and substantially enhances sample balance. Consequently, it offers an efficient and practical approximate optimization pathway for large-scale survey sampling and experimental design.

auxiliary variablesbalanced samplingcombinatorial optimization

This study addresses the limitation of existing self-training methods, where repetitive sampling induces homogeneous strategies and over-reliance on inherent model preferences, hindering performance on complex tasks. We propose a self-training principle centered on strategy diversity, introducing GROOT to construct hierarchical method trees that generate diverse reasoning trajectories, combined with Verbalized Sampling to optimize data construction. This work provides the first empirical evidence that strategy diversity is a stronger determinant of self-training efficacy than label correctness or teacher model scale. Experiments demonstrate that our approach significantly outperforms standard i.i.d. training on challenging benchmarks such as competitive programming. Notably, we show that smaller models trained on diversified error trajectories can surpass larger models optimized via conventional knowledge distillation.

Data constructionLarge language modelsSampling

This study addresses the limitation of existing machine learning methods, which prioritize predictive accuracy while neglecting design-unbiasedness—a critical requirement in official statistics and similar domains. The authors propose a general framework that does not rely on assumptions about the true data-generating model and, for the first time, integrates the known inclusion mechanisms from probability sampling designs into every stage of the learning pipeline: training sample selection, hyperparameter tuning, and performance evaluation. This integration guarantees design-unbiased prediction and classification over finite populations. The approach is compatible with popular algorithms such as k-nearest neighbors and random forests, establishes theoretical conditions under which design-unbiasedness is achieved, and provides practical algorithmic implementations alongside evaluation criteria.

algorithmic inferencefinite populationmachine learning

Hot Scholars

ZT

Zhaopeng Tu

Tech Lead @ Tencent Digital Human
Digital HumanAgentsLarge Language ModelsMachine Translation
PN

Preslav Nakov

Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Computational LinguisticsLarge Language ModelsFact-checkingFake News
PH

Pinjia He

Assistant Professor, The Chinese University of Hong Kong, Shenzhen
Software EngineeringAI4SESE4AIAIOps
JW

James Wang

Columbia University
Columbia University
WC

Wenhu Chen

Assistant Professor at University of Waterloo
Natural Language ProcessingArtificial IntelligenceDeep Learning