Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity

📅 2021-03-06
📈 Citations: 14
✨ Influential: 2
📄 PDF
🤖 AI Summary
This paper addresses the reinforcement bias and endogeneity arising from the dynamic coupling between data generation and policy evaluation in reinforcement learning (RL). We propose Instrumental Variable RL (IV-RL), the first RL paradigm explicitly designed to handle endogeneity via instrumental variables. By modeling policy iteration as a Markov process with instrument-dependent transitions, we develop an asymptotic statistical theory for IV-RL, derive an inferentially valid estimator of the optimal policy, and quantify how temporal dependence degrades inference accuracy. Our method integrates instrumental variable estimation, stochastic approximation theory, and Markov decision process (MDP) modeling. We rigorously establish strong consistency and asymptotic normality of the IV-RL estimator, enabling unbiased policy evaluation and principled confidence interval construction. The key contribution is breaking the conventional offline RL assumption of exogenous data: IV-RL provides the first statistically grounded, endogeneity-robust correction framework for dynamic decision-making under endogenous feedback.
📝 Abstract
In the standard data analysis framework, data is collected (once and for all), and then data analysis is carried out. However, with the advancement of digital technology, decision-makers constantly analyze past data and generate new data through their decisions. We model this as a Markov decision process and show that the dynamic interaction between data generation and data analysis leads to a new type of bias -- reinforcement bias -- that exacerbates the endogeneity problem in standard data analysis. We propose a class of instrument variable (IV)-based reinforcement learning (RL) algorithms to correct for the bias and establish their theoretical properties by incorporating them into a stochastic approximation (SA) framework. Our analysis accommodates iterate-dependent Markovian structures and, therefore, can be used to study RL algorithms with policy improvement. We also provide formulas for inference on optimal policies of the IV-RL algorithms. These formulas highlight how intertemporal dependencies of the Markovian environment affect the inference.
Problem

Research questions and friction points this paper is trying to address.

Markov Decision Process
Reinforcement Bias
Data Analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Instrumental Variables
Reinforcement Learning
Bias Mitigation in Markov Decision Processes
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
The University of Hong Kong | The Hong Kong University of Science and Technology
J
Jin Li
Faculty of Business and Economics, The University of Hong Kong, Pokfulam Road, Hong Kong SAR
Y
Ye Luo
Faculty of Business and Economics, The University of Hong Kong, Pokfulam Road, Hong Kong SAR
X
Xiaowei Zhang
Department of Industrial Engineering and Decision Analytics, The Hong Kong University of Science and Technology, Clear Water Bay, Hong Kong SAR