Score
Designs and implements procedures that draw multiple candidate outputs from a stochastic generator, score or rank those samples, and select the single best or a top-n subset for downstream use (sample-and-select / top-n selection). Analyzes and tunes the tradeoffs between number of samples, compute cost, selection criterion, and coverage of distinct modes to mitigate mode-selection failures and improve discovery of high-reward modes.
For single-problem optimization scenarios where extensive computational budgets permit multiple heuristic trials—e.g., testing algorithm variants, parameter configurations, initial solutions, or termination criteria—this work formally introduces the concept of “single-problem, multiple-attempt heuristic optimization.” We integrate algorithm selection, parameter adaptation, multi-start search, and dynamic resource allocation into a unified framework and propose a structured taxonomy. Our contributions are threefold: (i) we provide the first comprehensive survey of this underexplored direction; (ii) we unify cross-domain terminology and modeling paradigms; and (iii) we establish the first holistic taxonomy covering strategy design, decision mechanisms, and evaluation criteria. The framework offers theoretical foundations and practical guidelines for high-budget, single-problem optimization, significantly improving both efficiency and robustness of the multi-attempt process.
This paper addresses the low reliability of stochastic optimizer performance evaluation due to run-to-run variability. We propose a statistically grounded, adaptive experimental design method. First, we theoretically derive a lower bound on the minimum number of independent runs required to guarantee prescribed accuracy for key performance metrics—such as best objective value and convergence iteration count. Building upon this, we design an adaptive sampling algorithm that dynamically determines the requisite number of repetitions, ensuring termination only when both a user-specified confidence level (e.g., 95%) and absolute error tolerance are simultaneously satisfied—thereby avoiding premature stopping or unnecessary resource expenditure. The method integrates confidence interval estimation, hypothesis testing, and sequential sample-size determination, substantially enhancing reproducibility and statistical rigor in optimizer benchmarking and hyperparameter tuning. Empirical evaluation demonstrates that the approach consistently confines estimation error within the prescribed threshold while reducing redundant runs by over 30% on average.
This paper addresses the sequential selection problem under streaming heterogeneous inputs, where multi-source data collection and simulation execution must be coordinated under constrained budgets. Method: We propose the first synchronous budget allocation framework that constructs a performance estimator based on temporally aggregated heterogeneous simulation outputs and jointly optimizes data acquisition and simulation resource allocation. Theoretically, we establish asymptotic consistency and asymptotic normality of the estimator. Methodologically, we design a multi-stage stochastic optimization algorithm ensuring both statistical reliability and computational tractability. Results: Numerical experiments demonstrate that our approach significantly outperforms existing benchmarks in selection accuracy and resource utilization efficiency, providing a provably sound, computationally feasible, and practically deployable paradigm for real-time sequential decision-making under streaming heterogeneous environments.
This paper addresses the two-stage experimental practice—first selecting the best treatment, then estimating its effect—in multi-arm trials (e.g., clinical trials, A/B/n tests). We propose the first experimental design framework jointly optimizing both selection accuracy of the winning treatment and precision of its effect estimation. Methodologically, we unify these objectives into a single optimization criterion, leveraging Neyman’s allocation principle; a theory-driven algorithm determines optimal sample allocations across the control and all treatment arms. The framework provides finite-sample theoretical guarantees and achieves asymptotic optimality. Simulation studies and real-data experiments demonstrate that our approach significantly outperforms conventional balanced allocation and sequential designs in both winning-treatment identification accuracy and average treatment effect estimation precision.
This paper addresses the Stochastic Black-Box Optimization Selection (SBOS) problem: identifying, among multiple stochastic systems with continuous decision variables, the system whose optimal decision yields the best expected performance—without prior knowledge and under a finite sampling budget. We propose the first formal SBOS framework that jointly integrates intra-system optimization via stochastic gradient descent and inter-system comparison via sequential elimination, enabling synergistic optimization across both levels. We theoretically establish that the probability of incorrect selection converges exponentially with the sampling budget. Empirical evaluation across three real-world SBOS scenarios demonstrates that our method significantly reduces the misselection probability and maintains robust superiority across varying budget sizes and problem dimensions.
Scientific workflows often involve optimization objectives and evaluation criteria that are inherently uncertain and evolve with accumulating evidence, posing challenges for traditional Bayesian optimization methods. This work proposes the Generate-Select-Refine (GSR) framework, which uniquely integrates open-ended task discovery into the Bayesian optimization loop. Starting from user-provided seed tasks, GSR generates new tasks in a coarse-to-fine manner and employs a task acquisition function to orchestrate the optimization process, enabling alternating cycles of task discovery and refinement. The approach incurs only logarithmic regret overhead, transcending the limitations of single-task optimization. Empirical results demonstrate that GSR significantly outperforms existing large language model–based optimizers across diverse domains, including new product development, chemical process scale-up, algorithmic analysis, and patent repurposing.
Large language models are susceptible to selection bias in adaptive prompting and program search, leading to an overestimation of the winning candidate’s performance under real-world deployment. This work proposes the SIREN protocol, which enables unbiased performance inference for the full tuning-to-deployment pipeline under a fixed tuning budget by freezing the candidate set, decoupling selection and evaluation data, and incorporating an entry-wise Gaussian multiplier bootstrap. SIREN is the first method to simultaneously support accurate estimation of program-level performance curves on limited-budget grids and construct confidence intervals for both within-budget and cross-budget comparisons. Empirical results demonstrate that conventional winner-reporting practices exhibit substantial optimistic bias, whereas SIREN closely approximates the true evaluation target under finite-sample conditions, offering reliable guidance for deployment decisions.
This study addresses the challenge of efficiently identifying practically meaningful treatment effects under resource constraints and concurrent experimentation, where conventional resource allocation strategies—optimized to minimize mean squared error (MSE)—often prove suboptimal. The authors propose a novel framework that shifts the objective toward minimizing the worst-case Type II error (i.e., miss rate) by leveraging statistical power. They develop a variance inflation mechanism with a correction factor, tailored to scenarios where outcome standard deviations are either known or estimated from pilot data, and formulate optimization models under three distinct risk criteria. A fully data-driven Surrogate-S algorithm is introduced to implement the approach without requiring ground-truth variance information. Theoretical analysis demonstrates the potential inefficiency of MSE-oriented strategies in detection tasks, while numerical experiments show that the proposed method achieves near-optimal performance using only pilot-based variance estimates.
This work addresses the high cost of machine learning benchmarking by proposing a systematic framework to efficiently select small, representative subsets of datasets while preserving model ranking stability. The study presents the first comprehensive evaluation of various dataset selection strategies—including clustering, A/D-optimal experimental designs, random baselines, and a greedy farthest-first (FAFI) approach—on rank fidelity. It derives a theoretical upper bound on Spearman rank correlation error for FAFI and integrates bootstrap aggregation to yield statistically rigorous confidence intervals for comparing strategy performance. Empirical results demonstrate that as few as five datasets suffice to achieve 0.95 rank correlation in time series classification, significantly outperforming random selection in NLP tasks, though gains are limited in recommendation systems.
This work addresses the lack of statistical reliability guarantees in existing hyperparameter selection methods—such as grid search and Bayesian optimization—with respect to critical metrics like risk and safety. Building upon the learn-then-test (LTT) paradigm, the paper introduces a unified statistical framework that formulates hyperparameter selection as a multiple hypothesis testing problem, accommodating user-specified constraints on average risk, quantile risk, or information-theoretic measures. Leveraging tools from statistical inference—including p-values, e-values, and concentration inequalities—the method derives explicit, finite-sample bounds on error probabilities from first principles. This approach enables theoretically grounded validation and selection of hyperparameters, substantially enhancing the reliability and safety of AI systems in real-world deployment scenarios.