🤖 AI Summary
This study addresses the bias-variance trade-off inherent in single-step Bellman residuals for off-policy evaluation in reinforcement learning by proposing a mixed Bellman residual framework. Methodologically, the approach achieves a natural trade-off through a convex combination of first- and second-order residuals. Leveraging minimax optimization and sample splitting techniques, it constructs a data-dependent adaptive kernel critic along future feature directions, thereby overcoming the limitations of fixed function classes to jointly optimize approximation error and importance sampling variance. Simulation experiments and evaluations on MetaWorld tasks demonstrate that this mixed residual strategy significantly enhances the accuracy of value estimation in complex scenarios.
📝 Abstract
Evaluating a target policy using data generated by a different behavior policy remains a fundamental challenge in reinforcement learning. While most existing work relies on the standard one-step Bellman residual, we consider a convex combination of one-step and two-step residuals with a fixed mixing weight. In ideal settings, this mixed Bellman formulation can provide a natural bias--variance trade-off between approximation error under a restricted value-function class and the increased variance arising from multi-step importance weighting. To solve this mixed residual optimization, we adopt a minimax formulation involving a critic function. Unlike standard approaches that rely on a fixed functional class, we construct a data-dependent critic representation using predicted future feature directions which effectively induces a kernel adapted to the underlying transition dynamics. This allows the critic to focus on directions that are most relevant for the estimated Bellman error. To control overfitting, we use sample splitting to construct the critic and estimate the value function on separate data subsets. Simulation studies and MetaWorld tasks illustrate the effect of the mixing parameter and show that intermediate residual combinations can improve value estimation in challenging settings.