🤖 AI Summary
This work addresses a theoretical gap in existing frameworks that combine distributional reinforcement learning with maximum-entropy control, which lack contraction guarantees for the distributional soft Bellman operator under the Cramér geometry. The authors construct a cumulative distribution function–based soft Bellman operator in the policy evaluation phase and, for the first time, prove its √γ-contraction property under the Cramér distance. They further demonstrate that its unique fixed point arises from a unified first-moment condition rather than conventional boundedness assumptions. Additionally, the paper establishes an equivalent spectral-domain representation of this operator in a Hilbert space. These results provide a rigorous theoretical foundation for designing critics and analyzing evaluation errors in distributional soft policy iteration algorithms.
📝 Abstract
Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the policy evaluation step is governed by a distributional soft Bellman operator acting on entropy-regularised returns. Theoretical analysis of such an evaluation step requires a probability metric under which Bellman updates can be controlled, typically by showing that the operator contracts the distance between any two candidate return-distribution estimates. In this paper, we focus on the Cramér geometry, a cumulative distribution function (CDF)-based metric with an $L^2$ structure, and study whether the fixed-policy distributional soft Bellman operator has this contraction property and hence a unique fixed point under this metric. Working directly on an admissible CDF field domain, we formulate the CDF-level distributional soft Bellman operator, prove that it is a $\sqrtγ$-contraction, and obtain the corresponding unique fixed point together with convergent iterative policy evaluation. The CDF formulation also shows that this finite-Cramér-domain property follows from a uniform first-moment condition on the combined one-step reward entropy shift, rather than from separate uniform boundedness assumptions on the reward and entropy terms. We then transport the same evaluation problem to the spectral domain by conjugation, obtaining an equivalent Hilbert-space representation of the same decision process. Taken together, these results identify the Cramér-geometric Bellman fixed point associated with the policy-evaluation step of DSPI, providing a reference point for studying approximate critics, evaluation error, and critic-loss design in DSPI-style algorithms.