Score
Designs and evaluates algorithms and procedures that select predictive models under limited computational or labeling budgets. This includes building budgeted selection algorithms and mappings of accuracy versus compute/label trade-offs, recommending models for a given label budget, and accounting for tuning and prediction costs to identify regimes where particular choices are preferred.
In machine learning–guided design, a critical challenge is reliably selecting a design algorithm that satisfies user-specified success criteria—e.g., ensuring ≥10% of generated designs exceed a threshold on a target property. Method: We propose the first prediction-powered inference framework for algorithm selection. It integrates predictions from a learned model with held-out labeled data, employing density-ratio weighting for calibration. The method provides theoretical guarantees that the selected algorithm satisfies the desired success probability constraint with high confidence, supporting both known and estimable density ratios; it either returns the optimal algorithm or certifies infeasibility. Crucially, it jointly optimizes the predictive and generative models. Results: Evaluated on simulated protein and RNA design tasks, our approach significantly improves accuracy in identifying successful algorithms while strictly enforcing user-specified probabilistic constraints—empirically validating the practicality of its theoretical guarantees.
Existing data selection methods ignore computational budget constraints, leading to unstable performance across varying budgets—and sometimes even underperforming random selection. To address this, we propose Computation-Aware Data Selection (CADS), a budget-aware framework that treats computational budget as a core optimization variable and jointly models data selection and budget constraints via a bi-level optimization formulation. CADS introduces three key innovations: (i) a Hessian-free gradient estimator for efficient meta-gradient computation, (ii) probabilistic reparameterization of the selection policy, and (iii) an inner-loop penalty simplification strategy to accelerate convergence. Evaluated on diverse vision and language tasks, CADS achieves up to 14.42% absolute accuracy improvement over state-of-the-art baselines. To our knowledge, CADS is the first method to systematically incorporate computational budget into the data selection decision process, significantly enhancing both training efficiency and model generalization.
Per-instance algorithm selection (PIAS) takes advantage of complementarity between a set of algorithms by deciding which algorithm to run on a given instance. This decision is based on features of the instances, which, in the context of black-box optimization (BBO), require a part of the optimization budget to be computed. This raises two questions: (a) from which fraction of the budget spent on feature computation does PIAS become worth it for BBO, and (b) which fraction of the budget optimizes the tradeoff between feature accuracy and PIAS performance. To this end, we perform a broad study where PIAS with varying sampling budgets for feature computation is compared to the single best algorithm on a broad range of algorithm selection scenarios. These scenarios consist of two portfolio sizes, three problem sets, 4 dimensionalities, and 10 target budgets. We find that PIAS is viable for the majority of tested scenarios, even when as much as a quarter of the total budget is spent on feature computation. The tradeoff for the fraction of the budget spent on feature computation to maximize the benefit of PIAS is highly dependent on the specific AS scenario. Further, on average 20 percent of PIAS loss to the virtual best solver is explained by the budget spent on feature computation, highlighting the importance of properly accounting for the feature budget.
The algorithm selection and parameterization (ASP) domain lacks systematic surveys and empirical evaluations. Method: We propose the first standardized, meta-learning–driven ASP framework, built upon the largest ASP benchmark knowledge base to date—comprising 400 datasets and 4 million pre-trained models—and conduct large-scale comparative experiments across eight mainstream classifiers under diverse scenarios. Our evaluation integrates empirical performance modeling (EPM), feature engineering, and statistical significance testing to quantify accuracy, generalizability, and computational efficiency. Contribution/Results: This work delivers the first critical survey balancing methodological rigor with empirical breadth; reveals performance boundaries and applicability conditions of state-of-the-art ASP methods; and establishes a reproducible benchmark and practical selection guide for AutoML research and deployment.
This paper challenges the unverified implicit assumption in the predict-then-optimize paradigm that “higher prediction accuracy necessarily yields better downstream decisions,” particularly in multiclass classification settings. Method: We propose a controllable, interpretable multiclass prediction simulation framework that explicitly models error types and distributions, enabling systematic analysis of how classification errors affect decision quality in constrained optimization. Contribution/Results: Experiments on job scheduling and other combinatorial optimization tasks reveal a nonlinear relationship between prediction error and decision performance: improving prediction accuracy does not guarantee improved solution quality—and can even degrade decisions when error patterns shift. Our findings question the conventional coupling logic between prediction and optimization, providing theoretical foundations and practical guidance for designing, evaluating, and calibrating classifiers specifically tailored to decision objectives.
This work addresses the fundamental problem in graph neural networks and semi-supervised learning of selecting the most informative set of \( k \) vertices, given a graph and a labeling budget \( k \), to optimally support label inference across the entire graph. The authors propose a novel approximation algorithm grounded in combinatorial optimization and graph theory, which achieves—for the first time under standard budget constraints—a theoretical approximation ratio of \( \tilde{O}(\log^{1.5} n) \). This result overcomes key limitations of prior approaches that either relied on resource augmentation or lacked rigorous theoretical guarantees. The algorithm is both scalable and effective: its efficient heuristic variant handles large-scale graphs while maintaining high label prediction accuracy and consistently outperforming existing methods.
This study addresses the absence of a universally optimal optimizer in search-based software engineering and the instability of existing recommendations—such as NSGA-II—across varying evaluation budgets. Conducting large-scale experiments across 106 software engineering tasks, the authors evaluate 20 optimization algorithms under four distinct budget conditions. By integrating task characteristics with budget constraints, they perform clustering and performance analysis to propose a lightweight, budget-aware optimizer selection strategy. Relying solely on two easily obtainable task attributes and the evaluation budget, this approach uses a lookup-table mechanism to predict high-performing optimizers, achieving performance on par with or better than the ex post facto best choice on approximately 75% of held-out tasks. The complete experimental suite is publicly released to support reproducible research.
This work addresses the critical challenge of dynamically determining when to perform continual fine-tuning of foundation models on resource-constrained devices under limited computational budgets to maximize performance. The problem is formally cast, for the first time, as a constrained Markov decision process, where the state encompasses model performance, remaining compute budget, and the relevance of incoming data to the historical distribution. The authors propose an online decision-making strategy based on an Actor-Critic reinforcement learning framework; when fine-tuning gains are predictable, dynamic programming is also integrated for optimal scheduling. Experimental results demonstrate that the proposed approach improves accuracy by over 4% compared to strong baselines under identical budgets and achieves 97% of the performance of full-parameter fine-tuning using only 25% of the fine-tuning steps.
This study addresses the limitation of existing machine learning methods, which prioritize predictive accuracy while neglecting design-unbiasedness—a critical requirement in official statistics and similar domains. The authors propose a general framework that does not rely on assumptions about the true data-generating model and, for the first time, integrates the known inclusion mechanisms from probability sampling designs into every stage of the learning pipeline: training sample selection, hyperparameter tuning, and performance evaluation. This integration guarantees design-unbiased prediction and classification over finite populations. The approach is compatible with popular algorithms such as k-nearest neighbors and random forests, establishes theoretical conditions under which design-unbiasedness is achieved, and provides practical algorithmic implementations alongside evaluation criteria.