Offline Policy Evaluation via Mixed Bellman Residuals and Adaptive Critic Representations

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bias-variance trade-off inherent in single-step Bellman residuals for off-policy evaluation in reinforcement learning by proposing a mixed Bellman residual framework. Methodologically, the approach achieves a natural trade-off through a convex combination of first- and second-order residuals. Leveraging minimax optimization and sample splitting techniques, it constructs a data-dependent adaptive kernel critic along future feature directions, thereby overcoming the limitations of fixed function classes to jointly optimize approximation error and importance sampling variance. Simulation experiments and evaluations on MetaWorld tasks demonstrate that this mixed residual strategy significantly enhances the accuracy of value estimation in complex scenarios.
📝 Abstract
Evaluating a target policy using data generated by a different behavior policy remains a fundamental challenge in reinforcement learning. While most existing work relies on the standard one-step Bellman residual, we consider a convex combination of one-step and two-step residuals with a fixed mixing weight. In ideal settings, this mixed Bellman formulation can provide a natural bias--variance trade-off between approximation error under a restricted value-function class and the increased variance arising from multi-step importance weighting. To solve this mixed residual optimization, we adopt a minimax formulation involving a critic function. Unlike standard approaches that rely on a fixed functional class, we construct a data-dependent critic representation using predicted future feature directions which effectively induces a kernel adapted to the underlying transition dynamics. This allows the critic to focus on directions that are most relevant for the estimated Bellman error. To control overfitting, we use sample splitting to construct the critic and estimate the value function on separate data subsets. Simulation studies and MetaWorld tasks illustrate the effect of the mixing parameter and show that intermediate residual combinations can improve value estimation in challenging settings.
Problem

Research questions and friction points this paper is trying to address.

Offline Policy Evaluation
Reinforcement Learning
Bellman Residuals
Innovation

Methods, ideas, or system contributions that make the work stand out.

Offline Policy Evaluation
Mixed Bellman Residuals
Adaptive Critic Representations
Sample Splitting
Minimax Optimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Amitakshar Biswas
Department of Statistics, University of Illinois at Urbana-Champaign
Y
Yuhan Li
Department of Statistics, University of Illinois at Urbana-Champaign
Ruoqing Zhu
Ruoqing Zhu
University of Illinois Urbana-Champaign
Personalized MedicineReinforcement LearningRandom ForestsSurvival AnalysisDimension Reduction