Score
Designs and analyzes decision-theoretic stopping rules that determine whether to continue or halt an iterative process by computing expected utilities (for example expected value of information) and comparing the expected benefit of additional actions to their costs. This work includes building optimal-stopping formulations, EVPI-based criteria, cost- and budget-aware stopping thresholds, and marginal-rate rules that recommend stop/continue decisions while trading off effort, cost, and expected payoff.
This study addresses the suboptimal decision-making in document screening, where existing stopping strategies focus solely on recall while neglecting the actual costs and benefits of review tasks. To bridge this gap, the work introduces decision theory into the problem for the first time, deriving three adaptive stopping strategies grounded in the Expected Value of Perfect Information (EVPI). These strategies are integrated within a Technology-Assisted Review (TAR) framework to align stopping decisions with task-specific utility objectives. Empirical evaluations on the CLEF-IP patent dataset and medical systematic review corpora demonstrate that the proposed approach significantly improves net utility compared to prevailing stopping rules, achieving better alignment between screening outcomes and real-world review goals.
Early stopping in Bayesian optimization (BO) for expensive black-box function evaluations remains challenging due to the lack of principled, cost-aware criteria. Method: We propose an adaptive, hyperparameter-free stopping rule grounded in theoretical analysis. First, we establish a rigorous theoretical connection between state-of-the-art cost-aware acquisition functions—such as PBGI and log EI/cost—and cumulative evaluation cost, deriving tight, closed-form upper bounds without heuristic tuning. Second, we formulate a theoretically justified stopping criterion by integrating Pandora’s Box Gittins Index with unit-cost expected improvement. Results: Empirical evaluation on hyperparameter optimization and neural architecture search demonstrates that, when combined with PBGI, our rule achieves superior or competitive performance in cost-adjusted simple regret—outperforming or matching existing methods—while significantly improving evaluation efficiency and theoretical soundness.
This paper addresses how dynamic information design influences agents’ stopping times and action choices in optimal stopping problems, particularly when the principal lacks intertemporal commitment power to induce dynamically consistent stopping behavior. Method: We develop a unified framework that jointly models dynamic persuasion and optimal stopping, grounded in game theory, Bayesian updating, dynamic programming, and stopping-time theory. Contribution/Results: We prove that, for any agent preference structure, there exists an optimal information structure that guarantees dynamic consistency without commitment. This structure endogenously determines the optimal stopping time, state-contingent actions, and the path of information revelation. Our framework provides both a rigorous theoretical foundation and a computationally tractable design paradigm for information manipulation in sequential decision-making contexts—including algorithmic recommendation systems and regulatory interventions.
This paper addresses the pricing of variable annuities with a minimum guaranteed maturity benefit, modeling policyholder surrender behavior as an optimal stopping problem aimed at maximizing risk-neutral value—yielding an unbounded, time-varying, and discontinuous payoff structure. Using stochastic optimal control theory and variational inequality methods, complemented by PDE analysis and novel auxiliary value function constructions, we systematically characterize, for the first time, the nonemptiness and geometric structure of the surrender region—entirely determined by management fees and surrender charges. We derive three distinct representations of the value function, one of which is original to both actuarial science and American option literature. We establish a sufficient condition under which the optimal stopping time is necessarily delayed until maturity. Furthermore, we quantitatively elucidate the intrinsic mechanism through which fee structures govern the location of the surrender boundary and the intensity of early surrender incentives.
In Bayesian sequential trials, error rate evaluation relies on computationally expensive Monte Carlo simulations, hindering efficient optimization of sample size and decision thresholds. Method: This paper establishes, for the first time, analytical functional relationships between posterior and posterior predictive probabilities and sample size. Leveraging Bayesian decision theory and asymptotic analysis—combined with numerical fitting and error-rate inversion—the method enables precise error-rate assessment for any sample size using only two simulations, and rapidly identifies optimal design parameters. Contribution/Results: The approach drastically reduces computational cost while achieving error-rate control accuracy comparable to conventional simulation-based methods. In two real-world case studies, it attains exact error-rate calibration and accelerates design optimization by several orders of magnitude. This provides a scalable, verifiable, and highly efficient design paradigm for Bayesian adaptive trials.
This study addresses a central challenge in data-driven optimization: determining the optimal stopping time for data acquisition under parameter uncertainty by balancing sampling costs against information gains. The authors propose a Bayesian learning–based sequential data collection framework that explicitly models the trade-off between information gain and sampling cost, enabling a reward-driven adaptive stopping mechanism that jointly optimizes data acquisition and decision-making. Integrating Bayesian parameter updating, sequential decision theory, and stochastic programming, the approach formulates multiple stopping strategies within the newsvendor model. Numerical experiments demonstrate that the proposed strategies significantly reduce redundant sampling compared to fixed-budget and ex post optimal benchmarks while maintaining near-optimal decision performance.
This work addresses the lack of theoretically grounded stopping criteria in Bayesian optimization, which often leads to excessive function evaluations and no guarantees on solution quality. Focusing on the GP-UCB algorithm, the authors derive a tighter upper bound on instantaneous regret and leverage it to propose the first stopping criterion with $(\varepsilon,\delta)$-optimality guarantees. This criterion ensures that, upon termination, the returned solution is approximately optimal with high probability. Experimental results demonstrate that the proposed method significantly reduces the number of function evaluations while strictly maintaining solution quality, thereby enhancing optimization efficiency.
This study addresses the pervasive issue of dynamic inconsistency in classical statistical decision rules based on ex ante criteria—such as minimax regret—which often become irrational to follow once data are observed. The paper provides the first systematic characterization of this problem, develops a formal framework for analyzing dynamic consistency, and axiomatically introduces two novel optimality criteria. By integrating decision theory, axiomatic reasoning, and minimax regret analysis, the proposed criteria effectively prevent post-data deviations across a range of empirical settings, thereby ensuring that decision rules remain coherent before and after information updating. This approach rectifies the behavioral discrepancies inherent in existing methodologies and enhances their practical applicability.
This study investigates how decision-makers dynamically choose what to explore and when to stop before taking action. By reformulating the dynamic exploration–stopping problem as a static optimization with an information budget constraint, the authors introduce an information shadow price and establish that the optimal policy exhibits a concavified structure of stopping gains net of this shadow price. They precisely characterize, for the first time, the set of attainable joint distributions and uncover the critical role of time preference curvature in shaping exploration: convex preferences induce Poisson-like exploration, whereas concave preferences restrict stopping to specific time windows. The framework also elucidates the emergence of pure exploration phases and is unifiedly applied to real options, speed–accuracy trade-offs, and continuous-time exploration contests, clearly delineating the structural impact of time preferences and information costs on exploration behavior.
This work proposes a generalized multi-armed bandit and stopping problem framework grounded in behavioral preferences, departing from conventional modeling paradigms that rely on predefined states, rewards, and transition dynamics. Starting from the decision maker’s preferences over local temporal plans, the authors develop a generalized stopping representation through behavioral axioms and introduce a calendar-time cross-plan pricing mechanism. Under compact time constraints, this approach yields a rested-bandit model exhibiting index optimality. The key innovation lies in interpreting the index as the shadow price of advancing a local clock, thereby unifying—within a single preference-based framework—a diverse array of decision models, including expected utility, learning, robust, rank-dependent, Choquet, and Pandora’s box formulations. This constitutes the first theoretical foundation for index policies rooted entirely in preference theory.