Score
Designs and implements iterative active-set algorithms and screening rules that identify likely-active variables, prune irrelevant ones, and maintain feasible near‑optimal solutions to shrink optimization or estimation problems. Builds decreasing or coarse‑to‑fine active‑set hierarchies (including active‑set compositional models) that separate presence from positive subcomposition, handle exact zeros, and scale to high‑dimensional variable‑selection and model‑reduction workflows.
This paper investigates the worst-case complexity of active-set methods for maximizing convex quadratic functions subject to linear constraints. Addressing the open question raised at ICPO 2025—“Are constantly many objective evaluations sufficient?”—we construct, for the first time, a quadratic objective function via a recursively defined deformed product formulation, augmented with geometric projection techniques. This construction forces the algorithm to traverse exponentially many vertices along a parabolic boundary of the feasible region. Crucially, it avoids shortcuts induced by high-dimensional primal edges, thereby establishing the first unconditional exponential lower bound on the number of iterations that relies solely on convex quadratic polynomials. This result significantly improves upon prior lower bounds, which required super-logarithmic-degree polynomials. Our work provides a key advance in the complexity analysis of simplex-type methods under arbitrary pivot rules.
This paper addresses the problem of efficiently identifying the true hypothesis under limited samples for an active decision-maker. We propose a deterministic multi-stage hypothesis elimination algorithm that, at each stage, selects actions maximizing distributional distinguishability—measured by total variation distance (TVD)—to guide sequential sampling. Crucially, we couple hypothesis clustering with iterative pruning, enabling batch elimination of hypotheses within similarity-defined clusters. To our knowledge, this is the first work to establish a joint clustering–sampling optimization framework grounded in distributional similarity, where adaptive sampling is guided by a TVD-based minimax criterion. We derive necessary and sufficient conditions for asymptotic error convergence to zero. Theoretically, the algorithm achieves asymptotically optimal sample complexity, admits a provable error bound, and runs in polynomial time—substantially improving convergence efficiency in large hypothesis spaces.
This work addresses combinatorial optimization problems subject to unknown linear constraints. We propose an active learning framework for feasible region approximation that operates under a limited membership oracle query budget. Instead of explicitly modeling constraints, our method employs a mixed-integer quadratic programming (MIQP)-driven optimal sampling strategy, jointly leveraging support vector machines (SVMs) and a convex-optimization-inspired linear separation mechanism to efficiently identify the feasible boundary. Compared to conventional SVM-based margin sampling, our approach significantly improves both query efficiency and boundary estimation accuracy. Experiments on knapsack and university course scheduling problems demonstrate accelerated objective convergence—by 37%–62%—and an average 12.4% improvement in final solution quality. The core contribution lies in integrating MIQP into the active learning loop, yielding a theoretically interpretable and computationally tractable paradigm for constraint discovery.
This work addresses active learning in realistic settings where sampling is constrained to an accessible region, while prediction targets may lie outside this domain. To tackle this challenge, we propose a transductive active learning framework tailored to real-world prediction objectives, leveraging adaptive uncertainty minimization for efficient out-of-domain inference. Theoretically, we establish, for the first time under general regularity assumptions, the uniform consistency of the decision rule—guaranteeing convergence to the minimal uncertainty achievable by accessible data—providing strong and broadly applicable statistical guarantees. Methodologically, our approach integrates transductive learning, Bayesian optimization, and large language model (LLM) fine-tuning. Empirical evaluations demonstrate substantial improvements in sample efficiency on LLM active fine-tuning and safety-critical Bayesian optimization tasks, achieving state-of-the-art performance.
This paper addresses the level set estimation (LSE) problem in manufacturing quality control under input uncertainty—arising from equipment precision variations and operator deviations—and introduces, for the first time, *cost-dependent input-uncertainty LSE*: accurately identifying the region where product performance meets specifications under heterogeneous, multi-precision, and multi-cost inspection devices, while minimizing total inspection cost. We propose a Bayesian optimization–based active learning algorithm that explicitly accounts for input uncertainty, equipped with a cost-weighted acquisition function and supported by theoretical convergence guarantees. Experiments on synthetic benchmarks and real-world industrial datasets demonstrate that our method achieves significantly lower total inspection costs than state-of-the-art approaches, while maintaining high-accuracy level set estimates—thereby bridging theoretical rigor and practical engineering applicability.
This work addresses the computational challenges of solving large-scale convex mixed-integer quadratic programs (MIQPs), which arise in applications such as subset portfolio selection and become particularly difficult when the covariance matrix has a high condition number or weight constraints are tight. To tackle this, the authors propose DASH, a novel method that introduces a decreasing active-set hierarchy for dimensionality reduction in MIQP for the first time. DASH leverages active-set analysis to reduce problem dimensionality and integrates seamlessly with commercial solvers like Gurobi to enhance optimization efficiency. Experimental results demonstrate that DASH significantly outperforms Gurobi alone on a range of challenging portfolio instances, with solution quality improvements positively correlated with problem difficulty, thereby accelerating convergence and yielding higher-quality optimal solutions.
This work addresses the problem of efficiently approximating an unknown subadditive set function when only a subset of its values can be queried, aiming to minimize additive error and reduce optimization uncertainty caused by missing data. We present the first systematic analysis of the tightest possible lower and upper completion bounds for subadditive functions under additive error. Building on this theoretical foundation, we propose a prior-informed active querying strategy that dynamically selects the most informative subsets to query in both offline and online settings. Our theoretical analysis characterizes the completion error bounds across different function classes, and extensive experiments demonstrate that the proposed algorithm significantly outperforms baseline methods, effectively narrowing the gap in completion error.
本文提出一种基于置信区域的筛选框架,用于解决模拟系统可接受性问题,保证高概率筛选出所有或每个可接受系统,并支持并行化。
This work addresses the statistical dependence induced by data reuse in two-stage empirical risk minimization (ERM), a common issue in settings such as active learning. The authors propose an iterative ERM framework that avoids data splitting or oracle assumptions. Leveraging tools from high-dimensional statistics, convex optimization, and random matrix theory, they derive—for the first time—the precise asymptotic expression for test error under linear models and convex losses. Their analysis reveals a double-descent phenomenon in test error driven by data selection and characterizes the fundamental trade-off between labeling budget allocation and generalization performance. Applied to pool-based active learning, the theory accurately predicts the performance of second-stage estimators, highlighting the critical role of annotation strategies in shaping generalization error.
This work addresses geometric data pruning methods that rely on neighborhood similarity assumptions, which inherently introduce selection bias. Discarding this assumption, we reformulate unbiased subset selection from first principles as a variance minimization problem. Through a linear programming perspective, we construct high-dimensional polytopes and derive closed-form pairwise variance expressions, enabling an efficient vertex-walking algorithm for label-agnostic data pruning with strictly guaranteed statistical unbiasedness. Experiments across multiple benchmarks demonstrate that the proposed method outperforms uniform sampling and mainstream geometric approaches in accuracy, exhibiting particularly superior performance under small selection budgets while effectively reducing stochastic gradient descent (SGD) variance.