On the Computational Tractability of Robust Bandits

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The provided original TLDR contained only incomplete fragments and lacked specific research content. To demonstrate compliant academic expression, the following optimized template is generated based on a general AI research framework; please replace the bracketed information with your actual research details. This study addresses the [specific bottleneck] of existing methods in [target task] by proposing a novel framework based on [core algorithm/model]. By introducing [key innovative module or mechanism], the proposed method effectively resolves [specific technical challenge]. Experiments demonstrate that our approach achieves significant improvements over baseline models on [mainstream benchmark datasets], enhancing [core metric] by [X]% while reducing computational overhead. This work provides an efficient new paradigm for [relevant application domain], with code and models publicly released.
📝 Abstract
Learning when the environment does not belong to the learner's hypothesis class is typically handled using agnostic learning guarantees. However, for anything beyond supervised learning, agnostic guarantees are difficult to come by. Recently, imprecise bandits (Kosoy, 2025) (later renamed to robust bandits in Appel and Kosoy, 2025) were introduced as another approach to unrealizable learning in the bandits setting and a $Θ(\sqrt{T})$ regret learner was shown for a large class. However, no computational guarantees were provided. In this paper we identify a special case that admits a polynomial-time learner with $\tilde{O}(\sqrt{T})$ regret. We also show that several small generalizations of this special case are NP-hard thus indicating that the special case is at the boundary of what is tractable. It has been recently suggested (Kosoy, 2018) that computationally efficient learners for unrealizable learning problems are crucial for solving the AI alignment problem. This work is a small step in that direction.
Problem

Research questions and friction points this paper is trying to address.

Robust Bandits
Computational Tractability
Unrealizable Learning
Regret Minimization
AI Alignment
🔎 Similar Papers
No similar papers found.