🤖 AI Summary
This work proposes a generalized multi-armed bandit and stopping problem framework grounded in behavioral preferences, departing from conventional modeling paradigms that rely on predefined states, rewards, and transition dynamics. Starting from the decision maker’s preferences over local temporal plans, the authors develop a generalized stopping representation through behavioral axioms and introduce a calendar-time cross-plan pricing mechanism. Under compact time constraints, this approach yields a rested-bandit model exhibiting index optimality. The key innovation lies in interpreting the index as the shadow price of advancing a local clock, thereby unifying—within a single preference-based framework—a diverse array of decision models, including expected utility, learning, robust, rank-dependent, Choquet, and Pandora’s box formulations. This constitutes the first theoretical foundation for index policies rooted entirely in preference theory.
📝 Abstract
Bandit models typically begin with arms, states, rewards, and transition rules. This paper instead begins with preferences over stopped local contingent schedules: possible unfoldings of a responsibility, project, experiment, or opportunity in its own local time. Behavioral axioms on single schedules characterize a generalized stopping representation with current utility, local discounting, and a broad continuation aggregator. A common-tail compensation axiom then allows calendar time to be priced across schedules. Imposing a tight elapsed-calendar constraint generates a rested generalized bandit and yields index optimality: the index is the shadow price of advancing a local clock. Expected-utility, learning, robust, rank-dependent, Choquet, and Pandora models arise as special cases.