Score
Design, build, or analyze algorithms and selection mechanisms that decide which unlabeled data points, time segments, or queries to acquire labels or feedback for, covering pool-based sampling, acquisition functions, query strategies, uncertainty- or curiosity-driven criteria, and objective-guided selection. This includes methods to integrate human corrections and sparse signals, implement incremental or online updates, optimize sampling under time or attention budgets, and reduce total labeling effort for classification and other supervised learning tasks.
Addressing model generalization under high annotation costs and data scarcity, this work establishes a unified theoretical and methodological framework for low-resource learning. Theoretically, it introduces the first agnostic active sampling theory integrated into the PAC learning framework, rigorously characterizing the trade-off between generalization error and labeling complexity. Methodologically, it proposes four novel optimization mechanisms—gradient-aware sampling, meta-iterative optimization, manifold-geometric modeling, and large-model-driven augmentation—and synergistically integrates transfer learning, reinforcement-based feedback, and hierarchical structural modeling. Empirical evaluation demonstrates that the proposed approach significantly enhances model robustness and generalization performance under limited annotations. This work provides both an interpretable, scalable theoretical foundation and a practical paradigm for data-constrained AI systems.
This work addresses the insufficient informativeness of sample selection in active learning by proposing a novel acquisition criterion based on gradient discrepancy. Rooted in the generalization error bound established by Luo et al., this criterion is the first to integrate gradient discrepancy with theoretical generalization bounds, thereby providing a principled foundation for sample selection. The proposed method serves as a viable alternative to conventional uncertainty-based measures and naturally unifies label uncertainty and sample diversity, bridging two dominant active learning paradigms. Theoretical analysis corroborates its validity, and empirical results demonstrate that the criterion consistently outperforms existing active learning approaches across multiple benchmark datasets.
This work addresses the algorithm selection problem in black-box optimization (BBO) based on probing trajectories. We systematically evaluate 17 time-series classifiers on the BBOB benchmark, using three types of short-horizon performance trajectories as inputs. Employing leave-one-instance and leave-one-problem cross-validation, we demonstrate—for the first time—that classifier architecture critically determines trajectory-driven algorithm selection performance. Specifically, feature-engineering-based models (e.g., TSF) and interval-based models (e.g., ROCKET) significantly outperform end-to-end sequence models such as RNNs and Transformers. The best-performing classifier achieves an average accuracy over 12 percentage points higher than LSTM and InceptionTime. Our study provides the first reproducible and interpretable guideline for selecting time-series classifiers in trajectory-based algorithm selection, establishing a foundation for principled, data-driven BBO meta-algorithm design.
This study addresses the degradation of model performance caused by noisy labels in crowdsourced annotation. We propose a robust learning framework grounded in signal processing principles, modeling annotator behavior and label generation mechanisms. For the first time, we introduce tensor identifiability and nonnegative matrix factorization (NMF) into crowdsourced truth inference, systematically tackling label fusion and latent ground-truth estimation. Methodologically, we unify statistical modeling, tensor decomposition, NMF, and preference-based learning techniques (e.g., RLHF and DPO), designing multiple noise-robust algorithms with provable convergence and mechanistic interpretability. Evaluated on diverse crowdsourced benchmark datasets, our approach achieves substantial improvements: +12.7% average accuracy in ground-truth recovery and +8.3% average gain in downstream classification accuracy. This work establishes a novel theoretical foundation and provides practical tools for learning from noisy labels.
This work addresses combinatorial optimization problems subject to unknown linear constraints. We propose an active learning framework for feasible region approximation that operates under a limited membership oracle query budget. Instead of explicitly modeling constraints, our method employs a mixed-integer quadratic programming (MIQP)-driven optimal sampling strategy, jointly leveraging support vector machines (SVMs) and a convex-optimization-inspired linear separation mechanism to efficiently identify the feasible boundary. Compared to conventional SVM-based margin sampling, our approach significantly improves both query efficiency and boundary estimation accuracy. Experiments on knapsack and university course scheduling problems demonstrate accelerated objective convergence—by 37%–62%—and an average 12.4% improvement in final solution quality. The core contribution lies in integrating MIQP into the active learning loop, yielding a theoretically interpretable and computationally tractable paradigm for constraint discovery.
In noisy label learning, sample selection suffers from dual biases: data bias (imbalanced selection sets) and training bias (error accumulation). To address these issues, this paper proposes ITEM, a noise-tolerant expert model. ITEM is the first to jointly model and mitigate both biases in a unified framework. It introduces a lightweight multi-expert robust network architecture, integrated with a dual-weighted class-discriminative sampler and a hybrid mini-batch training strategy. Additionally, an error-robust optimization mechanism is incorporated to enhance generalization under label noise. Extensive experiments on multiple benchmark datasets with synthetic and real-world label noise demonstrate that ITEM consistently outperforms state-of-the-art methods, achieving average accuracy gains of 3–5% while reducing parameter count by over 20%. The source code is publicly available.
This work addresses the problem of selectivity estimation for linear queries—such as point and range queries—in dynamic databases. It introduces, for the first time, online learning theory to this setting, proposing an online estimation method based on histogram models and standard loss functions. The approach effectively adapts to time-varying data distributions and query workloads, delivering provably low regret in both static and dynamic environments. The core contribution lies in establishing tight upper and lower bounds on regret specifically for histogram-based linear queries, thereby providing the first formal online learning framework for dynamic selectivity estimation with theoretical performance guarantees.
This study addresses the challenge of quota optimization in small-scale two-sided matching markets—such as sorority recruitment—characterized by scarce data and multi-round interaction constraints. The authors propose a dynamic quota allocation framework that integrates machine learning with operations research: compatibility scores are predicted using random forests, and integer linear programming dynamically optimizes invitation quotas across rounds, followed by final matching via the deferred acceptance algorithm. A novel robust fallback mechanism is introduced to handle scenarios with weak predictive signals, alongside an interactive tool designed for coordinators. Evaluated on a dataset of only 282 samples, the method achieves a ROC-AUC of 0.5822, yields quota allocations highly consistent with human decisions, and produces final matches aligning with the actual 2025 recruitment outcomes at 96.4% individual-level consistency and 100% match feasibility.