Elicitation and Decision Geometry in Single-Index Bandits

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the two-armed contextual single-index bandit problem, where a shared unknown monotonic link function renders the optimal action dependent solely on the index direction. To this end, we propose the Natural Boundary Learning (NBL) algorithm, which leverages sequential Stein contrasts to directly learn the optimal decision boundary without estimating the reward function or the common link. Theoretically, by revealing an intrinsic connection between the decision stability coefficient and the geometry induced by convex potentials, we prove that NBL converges to the optimal boundary under local stability conditions. This analysis establishes an expected regret bound of O(log n). Numerical experiments corroborate the theoretically predicted stability mechanism and demonstrate the model's robustness against link function misspecification.
📝 Abstract
We study two-arm contextual bandits with arm-specific single indices and a shared unknown monotone link. Monotonicity makes the optimal action depend only on the contrast between the index directions, hence arm-specific reward functions need not be estimated. We introduce Natural Boundary Learning (NBL), a greedy procedure that uses a sequential Stein contrast to learn the optimal boundary directly, without estimating the reward functions or the common link. We characterize the local Riemannian dynamics of NBL through a decision stability coefficient balancing arm separation, link geometry, and the context distribution. We show that this stability is connected to the elicitation geometry of the underlying convex potential. Under local decision stability, NBL contracts toward the optimal boundary and achieves $O(\log n)$ expected regret. Numerical experiments illustrate the predicted stability regimes and compare NBL with a parametric greedy benchmark under link misspecification.
Problem

Research questions and friction points this paper is trying to address.

contextual bandits
single-index model
decision boundary
monotone link function
regret minimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Natural Boundary Learning
Single-Index Bandits
Decision Stability
Elicitation Geometry
Sequential Stein Contrast
🔎 Similar Papers
No similar papers found.