Nearly Optimal Fixed-Confidence Best-Arm Identification with 1-Bit Feedback

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the fixed-confidence best-arm identification problem under strict 1-bit feedback constraints within a distribution-free setting with finite variance. To this end, it proposes a time-uniform 1-bit mean estimation primitive based on randomized threshold queries, which is embedded into a successive elimination-style candidate-challenger algorithmic framework. Furthermore, a phase-adaptive clipping strategy is designed to dynamically match the current resolution, fundamentally leveraging a clipped tail integral identity to achieve efficient estimation. Theoretically, this work establishes an information-theoretic lower bound up to logarithmic penalties and demonstrates that the proposed algorithm attains near-optimal sample complexity. Specifically, its leading term differs from the theoretical lower bound only by lower-order logarithmic factors, thereby achieving information-theoretic near-optimality under such stringent communication constraints.
📝 Abstract
We study fixed-confidence best-arm identification under strict 1-bit feedback constraints. At each round, the learner selects an arm and a query set, and receives only a single bit indicating whether the sampled reward belongs to that set. We consider a distribution-free finite-variance setting with arm-wise localization, where direct empirical mean estimation is no longer available and clipping becomes unavoidable. We first formulate a time-uniform 1-bit mean-estimation primitive based on randomized threshold queries and a clipped tail-integral identity. We then embed this primitive into candidate-challenger best-arm identification algorithms. A fixed-clipping algorithm gives a simple anytime $(ε,δ)$-PAC guarantee, while a phased adaptive-clipping algorithm matches the clipping level to the current resolution and yields a gap-adaptive sample complexity. We also prove a $K$-arm worst-case information-theoretic lower bound showing that the logarithmic penalty caused by finite-variance 1-bit feedback is intrinsic. This bound matches the leading dependence of the phased algorithm up to lower-order $\log\log$ factors.
Problem

Research questions and friction points this paper is trying to address.

Best-Arm Identification
1-Bit Feedback
Fixed-Confidence
Finite-Variance
Sample Complexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Best-Arm Identification
1-Bit Feedback
Mean Estimation
Adaptive Clipping
Information-Theoretic Lower Bound
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Khang Luong
Hanoi University of Science and Technology
D
Dinh Thai Son
Hanoi University of Science and Technology
Hoang Ta
Hoang Ta
National University of Singapore
CombinatoricsQuantum information theoryOptimization
Hung The Tran
Hung The Tran
AI Center, VNPT Media
Machine LearningOptimizationReinforcement LearningLarge Language Models
T
Tuan Quang Dam
Hanoi University of Science and Technology