Deep Weighted Bellman Residual Minimization for $Q^*$ Estimation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of distribution shift, Q-value overestimation, and low sample efficiency in offline reinforcement learning by proposing a weighted Bellman residual minimization framework. The proposed method integrates expert demonstrations with behavioral data through density ratio estimation to approximate the optimal Q-function. Furthermore, it relaxes the conventional completeness assumption and establishes theoretical convergence guarantees linking density ratio estimation to excess risk bounds. Experimental results demonstrate that this framework significantly improves numerical performance and policy generalization capability. Overall, this work provides both theoretical foundations and methodological guidance for the efficient utilization of expert data in offline reinforcement learning settings.
📝 Abstract
Off-policy evaluation is a foundational component of offline reinforcement learning, aiming to assess and optimize policy performance using pre-collected datasets. However, such datasets often suffer from pronounced challenges, including distribution shift, $Q$-value overestimation, and low sample utilization efficiency. To address these issues, this paper introduces a weighted Bellman residual minimization framework that incorporates density ratio weighting by effectively integrating expert demonstrations with behavioral data. The proposed weighting scheme departs from the conventional completeness assumption commonly imposed in the theoretical analysis of deep reinforcement learning. We establish a sharp convergence rate for density ratio estimation and derive the convergence rate for the excess risk of resulting deep $Q^*$ estimator. Extensive empirical evaluations demonstrate that, compared to existing methods, our method achieves significant improvements in numerical performance and policy generalization, providing specific guidance for the rational utilization of expert demonstrations.
Problem

Research questions and friction points this paper is trying to address.

off-policy evaluation
offline reinforcement learning
distribution shift
Q-value overestimation
sample efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Off-policy evaluation
Weighted Bellman residual minimization
Density ratio weighting
Expert demonstrations
Convergence rate
L
Lican Kang
Institute for Math and AI, Wuhan University, Wuhan, 430072, China; Hubei Key Laboratory of Computational Science, Wuhan University, Wuhan, 430072, China
J
Jerry Zhijian Yang
Institute for Math and AI, Wuhan University, Wuhan, 430072, China; School of Mathematics and Statistics, Wuhan University, Wuhan, 430072, China; Hubei Key Laboratory of Computational Science, Wuhan University, Wuhan, 430072, China
Cheng Yuan
Cheng Yuan
Associate Professor, School of Mathematics and Statistics, Central China Normal University
Computational PhysicsDeep Learning
C
Chen Zhong
School of Mathematics and Statistics, Wuhan University, Wuhan, 430072, China