🤖 AI Summary
This work investigates the formal origins of the Bellman equation and its conditions of applicability in dynamic programming and sequential decision-making. By analyzing the consistency among dynamic decomposition, return recursiveness, and uncertainty aggregation, the paper constructs a generalized Bellman recursion framework. It unifies, for the first time, three dual relationships—probability–return, return–aggregation, and aggregation–probability—within a single mathematical structure, thereby revealing the essential source of Bellman-type formulations. Leveraging analyses of sufficient statistics, recursive return modeling, and compatibility of uncertainty aggregation, the theory provides principled pathways—through state augmentation or structural transformation—for solving problems that violate classical Bellman assumptions, thereby integrating diverse methodologies from reinforcement learning, control theory, and decision theory.
📝 Abstract
What gives the Bellman equation its form? We show that the recursive properties of optimal value functions follow from three conditions: that the dynamics decomposes through sufficient statistics, that the return decomposes recursively, and that the aggregation of uncertainty is compatible with both. When all three conditions hold on a common state, the Bellman equation arises from their mutual consistency; when one fails, tractability can often be recovered by augmenting the state or by deforming return or dynamics. The same conditions are shown to give rise to three dualities: one between probability and return, one between return and aggregation, and one between aggregation and probability. Our framework reveals these dualities as arising from a single construction, unifying methods developed separately across reinforcement learning, control, and decision theory.