Score
Design and analyze algorithms, policies, and stopping rules that sequentially identify the best alternative(s) or top-k subset under uncertainty by allocating samples or comparisons (including adaptive pairwise and active top-k selection) to minimize measurements while meeting error criteria. Build acquisition and decision procedures—such as information-aware or Bayesian posterior-guided sampling, pure-exploration and best-arm identification routines, and top-k-aware stopping rules—that prioritize ambiguous items near the threshold, allow multiple correct answers or temporarily unanswerable estimates, and provide fixed-confidence or fixed-precision guarantees on selection accuracy.
This paper addresses complex pure-exploration objectives beyond best-arm identification—such as threshold testing and ε-optimal arm identification—by establishing the first duality-based minimax-optimal sampling allocation framework. It provides the first necessary and sufficient conditions for optimal sampling allocation in pure exploration; generalizes the top-two paradigm to arbitrary pure-exploration problems; and proposes a hyperparameter-free, information-directed selection rule driven by KL divergence and entropy. The rule is rigorously proven to achieve asymptotic optimality in Gaussian settings and resolves the long-standing open problem of asymptotic optimality for top-two Thompson sampling. Experiments demonstrate substantial improvements in sampling efficiency across Gaussian best-arm identification, threshold-bandwidth testing, and ε-optimal arm identification, consistently outperforming state-of-the-art methods.
In sequential Bayesian experimental design, pre-specifying a fixed number of experiments fails to accommodate dynamic real-world requirements, necessitating a principled solution to the fundamental “optimal stopping” problem. This paper pioneers the integration of optimal stopping theory into this framework by formulating a Markov decision process that jointly optimizes stopping policies and experimental designs; we prove that the optimal stopping rule balances immediate reward against the expected value of continuing. To address policy circular dependency during training, we propose a curriculum learning strategy to enhance convergence stability. Our method unifies Bayesian inference, policy gradient optimization, and curriculum learning. Empirical evaluation on linear-Gaussian benchmark tasks and contaminant source localization demonstrates significant improvements in estimation accuracy and sampling efficiency over baseline methods (e.g., fixed-threshold rules), especially under strong sequential dependencies.
In Bayesian sequential trials, error rate evaluation relies on computationally expensive Monte Carlo simulations, hindering efficient optimization of sample size and decision thresholds. Method: This paper establishes, for the first time, analytical functional relationships between posterior and posterior predictive probabilities and sample size. Leveraging Bayesian decision theory and asymptotic analysis—combined with numerical fitting and error-rate inversion—the method enables precise error-rate assessment for any sample size using only two simulations, and rapidly identifies optimal design parameters. Contribution/Results: The approach drastically reduces computational cost while achieving error-rate control accuracy comparable to conventional simulation-based methods. In two real-world case studies, it attains exact error-rate calibration and accelerates design optimization by several orders of magnitude. This provides a scalable, verifiable, and highly efficient design paradigm for Bayesian adaptive trials.
This paper addresses sequential experimental design under high sampling costs: how to adaptively allocate units to two treatments and dynamically determine stopping time to maximize expected welfare net of sampling costs (i.e., minimize worst-case regret). Methodologically, it proposes a unified stopping rule based on the product of the estimated average treatment effect and sample size exceeding a cost-dependent threshold, coupled with Neyman allocation for unit assignment. Theoretically, it establishes—first rigorously—that Neyman allocation is minimax-optimal for sampling, and characterizes the optimal stopping time via this threshold-crossing condition. The policy remains minimax-optimal under both parametric and nonparametric settings and in the small-cost asymptotic regime. It unifies Wald’s sequential hypothesis testing and best-arm identification as special cases. The solution admits a closed-form analytic expression and retains robust optimality even in the zero-cost limit.
Existing fixed-confidence best-arm identification (BAI) methods either require solving an optimization problem per round or enforce uniform exploration, limiting adaptability to non-Gaussian reward distributions. Method: We propose a novel Bayesian BAI strategy that integrates Thompson sampling with an optimal challenger rule—marking the first natural incorporation of Thompson sampling into the fixed-confidence BAI framework. It eliminates real-time optimization and minimum exploration constraints. Contribution/Results: We establish a β-optimality analysis framework: proving asymptotic optimality for the two-arm case and providing approximate optimality guarantees for $K geq 3$. Empirically, our method achieves sample complexity competitive with asymptotically optimal algorithms while significantly reducing computational overhead. The core innovation lies in introducing a new Bayesian BAI paradigm that jointly ensures theoretical rigor—via provable near-optimal sample complexity—and computational efficiency—through closed-form posterior updates and no per-round optimization.
This work addresses the problem of reliably identifying the top-$k$ items with minimal expected sample complexity under a fixed confidence level, by actively selecting pairwise comparisons. Focusing on a noisy pairwise comparison setting grounded in a latent utility model, the study presents the first asymptotically optimal algorithm, which integrates saddle-point optimization, primal-dual online learning, and an adaptive comparison allocation strategy. As the error probability vanishes, the algorithm achieves the information-theoretic lower bound. Theoretical analysis unveils the saddle-point structure underlying this fundamental limit, and the proposed method attains asymptotic optimality in sample efficiency, substantially advancing the theoretical foundations of active ranking learning.
This study addresses a central challenge in data-driven optimization: determining the optimal stopping time for data acquisition under parameter uncertainty by balancing sampling costs against information gains. The authors propose a Bayesian learning–based sequential data collection framework that explicitly models the trade-off between information gain and sampling cost, enabling a reward-driven adaptive stopping mechanism that jointly optimizes data acquisition and decision-making. Integrating Bayesian parameter updating, sequential decision theory, and stochastic programming, the approach formulates multiple stopping strategies within the newsvendor model. Numerical experiments demonstrate that the proposed strategies significantly reduce redundant sampling compared to fixed-budget and ex post optimal benchmarks while maintaining near-optimal decision performance.
This work addresses the lack of theoretically grounded stopping criteria in Bayesian optimization, which often leads to excessive function evaluations and no guarantees on solution quality. Focusing on the GP-UCB algorithm, the authors derive a tighter upper bound on instantaneous regret and leverage it to propose the first stopping criterion with $(\varepsilon,\delta)$-optimality guarantees. This criterion ensures that, upon termination, the returned solution is approximately optimal with high probability. Experimental results demonstrate that the proposed method significantly reduces the number of function evaluations while strictly maintaining solution quality, thereby enhancing optimization efficiency.
This work addresses the limitations of conventional sensor selection methods that focus exclusively on identifying a single optimal hypothesis (top-1), which is insufficient for target localization tasks requiring multiple high-probability candidate nodes. To overcome this, the authors propose a set-based decision rule grounded in top-p hypothesis coverage, integrating geometric awareness within a sequential hypothesis testing framework to design a novel sensor selection algorithm. By replacing the traditional top-1 criterion with a top-p performance metric, the approach better aligns with practical demands for generating reliable candidate sets. Experimental results on a real-world platform demonstrate that the proposed method significantly improves coverage performance of the selected sensor node lists, outperforming existing top-1 strategies.