Score
Designs and implements batch-sampling algorithms that construct training mini-batches with controlled class proportions (e.g., class-balanced or batch-balanced) by selecting samples from current data and optionally from stored or high-confidence historical examples; these procedures reduce bias toward dominant classes and stabilize model updates during training or online adaptation.
Class imbalance severely degrades model discrimination for minority classes, critically hindering deployment in high-stakes domains such as healthcare and finance. This paper systematically surveys over one hundred imbalance mitigation strategies, introducing the first unified taxonomy that integrates generative approaches (e.g., GANs, VAEs) with classical resampling techniques—including SMOTE, neighborhood density estimation, and adaptive threshold-based resampling. We further propose a multidimensional evaluation framework and practical deployment guidelines tailored to real-world constraints. Empirical validation across diverse benchmark tasks demonstrates that the surveyed methods improve minority-class F1-score by 12–35%. Crucially, we identify a novel pathway for jointly optimizing interpretability and generalization—bridging theoretical advances with engineering feasibility. This work provides a comprehensive, actionable foundation for both advancing imbalance learning theory and enabling robust, trustworthy deployment in critical applications.
This work proposes a self-balancing sequential sampling method tailored for applications such as auditing, scheduling, and sampling, where both unpredictability and distributional convergence are critical. By adaptively adjusting selection probabilities, the method achieves optimal $O(n^{-1})$ convergence of the empirical distribution to the target distribution while preserving maximal sample unpredictability—substantially improving upon the $O(n^{-1/2})$ rate of conventional i.i.d. sampling. Theoretical analysis reveals that this approach uniquely solves an entropy-regularized optimization problem and exhibits a diffusion limit linked to the Ornstein-Uhlenbeck process. Empirical results confirm its effectiveness in minimizing repeated selections and coverage gaps, all while rigorously maintaining control over unpredictability.
研究针对长尾图像分类问题,通过比较四种mini-batch采样策略,发现渐进平衡采样在提高尾部类别精度方面优于其他方法。
In noisy label learning, sample selection suffers from dual biases: data bias (imbalanced selection sets) and training bias (error accumulation). To address these issues, this paper proposes ITEM, a noise-tolerant expert model. ITEM is the first to jointly model and mitigate both biases in a unified framework. It introduces a lightweight multi-expert robust network architecture, integrated with a dual-weighted class-discriminative sampler and a hybrid mini-batch training strategy. Additionally, an error-robust optimization mechanism is incorporated to enhance generalization under label noise. Extensive experiments on multiple benchmark datasets with synthetic and real-world label noise demonstrate that ITEM consistently outperforms state-of-the-art methods, achieving average accuracy gains of 3–5% while reducing parameter count by over 20%. The source code is publicly available.
This work addresses the problem of deriving generalization error upper bounds for batch learning algorithms under mixing stochastic processes (i.e., dependent data), without imposing any stability assumptions on the batch learner. The method introduces a novel analytical framework based on Online-to-Batch conversion: stability requirements are shifted to an associated online learner, and a new notion of online algorithm stability—defined via the first-order Wasserstein distance—is proposed for the first time. It is shown that the Exponentially Weighted Average (EWA) algorithm satisfies this stability condition. Consequently, the framework yields both expectation- and high-probability generalization bounds applicable to *any* batch learning algorithm; under i.i.d. assumptions, the bounds reduce to classical forms up to correction terms governed by the mixing decay rate. The resulting bounds are explicitly computable, substantially broadening the applicability and practical utility of learning theory under data dependence.
To address the challenge of jointly maximizing predictive mean, uncertainty, and minimizing intra-batch redundancy in batch Bayesian optimization (BO), this paper proposes Thompson Sampling with Regret-to-Sigma Ratio (TS-RSR). TS-RSR is the first method to incorporate the regret-to-sigma ratio into the batch acquisition objective, leveraging a Thompson sampling approximation to explicitly balance exploration and exploitation across batch points. We provide theoretical guarantees of convergence. By integrating Gaussian process modeling with an efficient batch sampling scheme, TS-RSR significantly reduces intra-batch redundancy. Extensive experiments on synthetic benchmarks and real-world tasks demonstrate that TS-RSR consistently outperforms state-of-the-art batch BO methods, achieving new state-of-the-art performance.
This study addresses the combinatorial optimization problem of identifying a balanced sampling design from a large population under a fixed inclusion probability, such that the weighted estimator of an auxiliary variable closely approximates the known population total—a task of exponential complexity. To tackle this challenge, the authors propose a heuristic approach based on genetic algorithms, which iteratively refines the sampling scheme by integrating minimum support designs with candidate samples exhibiting high balance. This method overcomes the limitations of the traditional cube method in achieving balance and substantially enhances sample balance. Consequently, it offers an efficient and practical approximate optimization pathway for large-scale survey sampling and experimental design.
This study addresses the limitation of existing self-training methods, where repetitive sampling induces homogeneous strategies and over-reliance on inherent model preferences, hindering performance on complex tasks. We propose a self-training principle centered on strategy diversity, introducing GROOT to construct hierarchical method trees that generate diverse reasoning trajectories, combined with Verbalized Sampling to optimize data construction. This work provides the first empirical evidence that strategy diversity is a stronger determinant of self-training efficacy than label correctness or teacher model scale. Experiments demonstrate that our approach significantly outperforms standard i.i.d. training on challenging benchmarks such as competitive programming. Notably, we show that smaller models trained on diversified error trajectories can surpass larger models optimized via conventional knowledge distillation.
本文提出了一种基于Stein变分梯度下降的抽样框架,将完全顺序设计方法转换为批量顺序设计方法,以解决实验设计中的批量选择问题。
该研究通过将训练子样本视为第二阶段抽样,解决了模型辅助估计中的不确定性量化问题,并提出了两种方差估计方法。
This study addresses the limitation of existing machine learning methods, which prioritize predictive accuracy while neglecting design-unbiasedness—a critical requirement in official statistics and similar domains. The authors propose a general framework that does not rely on assumptions about the true data-generating model and, for the first time, integrates the known inclusion mechanisms from probability sampling designs into every stage of the learning pipeline: training sample selection, hyperparameter tuning, and performance evaluation. This integration guarantees design-unbiased prediction and classification over finite populations. The approach is compatible with popular algorithms such as k-nearest neighbors and random forests, establishes theoretical conditions under which design-unbiasedness is achieved, and provides practical algorithmic implementations alongside evaluation criteria.