Dual-Directed Algorithm Design for Efficient Pure Exploration

📅 2023-10-30
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses complex pure-exploration objectives beyond best-arm identification—such as threshold testing and ε-optimal arm identification—by establishing the first duality-based minimax-optimal sampling allocation framework. It provides the first necessary and sufficient conditions for optimal sampling allocation in pure exploration; generalizes the top-two paradigm to arbitrary pure-exploration problems; and proposes a hyperparameter-free, information-directed selection rule driven by KL divergence and entropy. The rule is rigorously proven to achieve asymptotic optimality in Gaussian settings and resolves the long-standing open problem of asymptotic optimality for top-two Thompson sampling. Experiments demonstrate substantial improvements in sampling efficiency across Gaussian best-arm identification, threshold-bandwidth testing, and ε-optimal arm identification, consistently outperforming state-of-the-art methods.
📝 Abstract
While experimental design often focuses on selecting the single best alternative from a finite set (e.g., in ranking and selection or best-arm identification), many pure-exploration problems pursue richer goals. Given a specific goal, adaptive experimentation aims to achieve it by strategically allocating sampling effort, with the underlying sample complexity characterized by a maximin optimization problem. By introducing dual variables, we derive necessary and sufficient conditions for an optimal allocation, yielding a unified algorithm design principle that extends the top-two approach beyond best-arm identification. This principle gives rise to Information-Directed Selection, a hyperparameter-free rule that dynamically evaluates and chooses among candidates based on their current informational value. We prove that, when combined with Information-Directed Selection, top-two Thompson sampling attains asymptotic optimality for Gaussian best-arm identification, resolving a notable open question in the pure-exploration literature. Furthermore, our framework produces asymptotically optimal algorithms for pure-exploration thresholding bandits and $varepsilon$-best-arm identification (i.e., ranking and selection with probability-of-good-selection guarantees), and more generally establishes a recipe for adapting Thompson sampling across a broad class of pure-exploration problems. Extensive numerical experiments highlight the efficiency of our proposed algorithms compared to existing methods.
Problem

Research questions and friction points this paper is trying to address.

Develops optimal adaptive experimentation for pure-exploration goals
Extends top-two approach beyond best-arm identification
Resolves asymptotic optimality in Gaussian best-arm identification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual variables optimize sampling allocation
Hyperparameter-free Information-Directed Selection rule
Asymptotically optimal top-two Thompson sampling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Columbia University | The Hong Kong University of Science and Technology
C
Chao Qin
Columbia Business School, Columbia University
W
Wei You
Department of IEDA, The Hong Kong University of Science and Technology