🤖 AI Summary
This study addresses the lack of an axiomatic foundation for entropy-regularized policies—such as the Boltzmann policy—and resolves their apparent conflict with the independence axiom in Markov decision processes. By distinguishing environmental randomness (chance) from decisional randomness (choice), the authors impose the Independence of Irrelevant Alternatives (IIA) and monotonicity axioms at choice nodes. Combining these with von Neumann–Morgenstern expected utility theory and a generalized discounting mechanism, they uniquely derive the Boltzmann policy and the soft Bellman equation from a normative axiomatic system. This work establishes a rigorous axiomatic basis for entropy regularization, revealing that the fundamental distinction between soft and hard Bellman equations lies in whether the agent values its own capacity to choose. It further proves convergence under return monotonicity and generalized discounting, unifying previously independent derivations from economics and information theory.
📝 Abstract
The softmax policy $π(a \mid s) \propto \exp(βQ(s,a))$ is the default model of stochastic choice in reinforcement learning (RL). Various justifications based on robustness, exploration, and optimization have been offered in the RL literature, but none uniquely derives the softmax form from first principles. This leaves a basic tension unresolved: the entropy bonus in the soft Bellman equation violates the Independence axiom that underwrites the Markov decision process (MDP) reward structure. We dissolve this tension by distinguishing two kinds of randomness: chance and choice. By restricting von Neumann-Morgenstern (VNM) Independence to environmental lotteries over base prospects, we show that imposing independence of irrelevant alternatives (IIA) and monotonicity on the policy and value functions at choice nodes uniquely determines the Boltzmann policy, the entropy-regularized representation, and the soft Bellman equation. The choice between the soft and hard Bellman equations thus reduces to a design decision: whether the agent values its own ability to choose. We develop RL-specific consequences, including return monotonicity and convergence under generalized discounting, and synthesize the independent lines from economics and information theory that arrive at the same structure, offering a normative assessment of when IIA is appropriate for agent design.