🤖 AI Summary
This study addresses the fixed-confidence best-arm identification problem under strict 1-bit feedback constraints within a distribution-free setting with finite variance. To this end, it proposes a time-uniform 1-bit mean estimation primitive based on randomized threshold queries, which is embedded into a successive elimination-style candidate-challenger algorithmic framework. Furthermore, a phase-adaptive clipping strategy is designed to dynamically match the current resolution, fundamentally leveraging a clipped tail integral identity to achieve efficient estimation. Theoretically, this work establishes an information-theoretic lower bound up to logarithmic penalties and demonstrates that the proposed algorithm attains near-optimal sample complexity. Specifically, its leading term differs from the theoretical lower bound only by lower-order logarithmic factors, thereby achieving information-theoretic near-optimality under such stringent communication constraints.
📝 Abstract
We study fixed-confidence best-arm identification under strict 1-bit feedback constraints. At each round, the learner selects an arm and a query set, and receives only a single bit indicating whether the sampled reward belongs to that set. We consider a distribution-free finite-variance setting with arm-wise localization, where direct empirical mean estimation is no longer available and clipping becomes unavoidable. We first formulate a time-uniform 1-bit mean-estimation primitive based on randomized threshold queries and a clipped tail-integral identity. We then embed this primitive into candidate-challenger best-arm identification algorithms. A fixed-clipping algorithm gives a simple anytime $(ε,δ)$-PAC guarantee, while a phased adaptive-clipping algorithm matches the clipping level to the current resolution and yields a gap-adaptive sample complexity. We also prove a $K$-arm worst-case information-theoretic lower bound showing that the logarithmic penalty caused by finite-variance 1-bit feedback is intrinsic. This bound matches the leading dependence of the phased algorithm up to lower-order $\log\log$ factors.