🤖 AI Summary
This paper addresses robust online decision-making under structured observations. We propose a novel framework based on multi-distribution probabilistic models: for each action, the outcome distribution belongs to a convex set, and the environment may adversarially select the true distribution from this set in a non-stationary, history-dependent manner. This setting relaxes the strong realizability assumptions prevalent in classical reinforcement learning and multi-armed bandits, enhancing practical relevance. We establish, for the first time, a theoretical foundation for robust multi-distribution modeling and introduce the notion of “power-law learnability,” providing its complete characterization. Building upon this, we derive tight regret upper and lower bounds for robust linear bandits and tabular robust online reinforcement learning—substantially improving upon prior results. Our analysis identifies power-law learnability as the fundamental scaling law governing learnability in this framework.
📝 Abstract
We propose a framework which generalizes"decision making with structured observations"by allowing robust (i.e. multivalued) models. In this framework, each model associates each decision with a convex set of probability distributions over outcomes. Nature can choose distributions out of this set in an arbitrary (adversarial) manner, that can be nonoblivious and depend on past history. The resulting framework offers much greater generality than classical bandits and reinforcement learning, since the realizability assumption becomes much weaker and more realistic. We then derive a theory of regret bounds for this framework. Although our lower and upper bounds are not tight, they are sufficient to fully characterize power-law learnability. We demonstrate this theory in two special cases: robust linear bandits and tabular robust online reinforcement learning. In both cases, we derive regret bounds that improve state-of-the-art (except that we do not address computational efficiency).