Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation

📅 2025-05-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper investigates a counterintuitive phenomenon in off-policy evaluation (OPE): even when the true behavior policy is Markovian, employing history-dependent (non-Markovian) estimators in importance sampling can reduce mean squared error (MSE). We derive the first rigorous bias–variance decomposition for importance sampling estimators in OPE and prove that history-dependent modeling substantially reduces asymptotic variance—monotonically decreasing with increasing history length. Methodologically, we integrate sequential importance sampling, marginalized weights, doubly robust estimation, and both parametric and nonparametric behavior policy modeling. Theoretically and empirically, we demonstrate that history-dependent estimators effectively balance bias and variance in finite-sample regimes, yielding significant gains in OPE accuracy. Our work establishes a novel paradigm for behavior policy modeling in OPE, challenging the conventional assumption that Markovian approximations are always optimal.

Technology Category

Machine Learning: Online Learning & BanditsSearch and Optimization: Sampling/Simulation-based SearchReasoning under Uncertainty: Sequential Decision Making

Application Category

User Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating successSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Human-perceived consequences of algorithmic deployment on the web
📝 Abstract
This paper studies off-policy evaluation (OPE) in reinforcement learning with a focus on behavior policy estimation for importance sampling. Prior work has shown empirically that estimating a history-dependent behavior policy can lead to lower mean squared error (MSE) even when the true behavior policy is Markovian. However, the question of why the use of history should lower MSE remains open. In this paper, we theoretically demystify this paradox by deriving a bias-variance decomposition of the MSE of ordinary importance sampling (IS) estimators, demonstrating that history-dependent behavior policy estimation decreases their asymptotic variances while increasing their finite-sample biases. Additionally, as the estimated behavior policy conditions on a longer history, we show a consistent decrease in variance. We extend these findings to a range of other OPE estimators, including the sequential IS estimator, the doubly robust estimator and the marginalized IS estimator, with the behavior policy estimated either parametrically or non-parametrically.
Problem

Research questions and friction points this paper is trying to address.

Explains why history-dependent behavior policy reduces MSE in OPE
Analyzes bias-variance trade-off in importance sampling estimators
Extends findings to various OPE estimators and estimation methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

History-dependent behavior policy reduces variance
Bias-variance tradeoff in importance sampling
Extends findings to multiple OPE estimators
🔎 Similar Papers
No similar papers found.
Hongyi Zhou
Hongyi Zhou
Karlsruhe Institute of Technology
reinforcement learningimitation learningrobotics
J
Josiah P. Hanna
Computer Sciences Department, University of Wisconsin – Madison, Madison, WI, USA
J
Jin Zhu
Y
Ying Yang
Department of Mathematical Science, Tsinghua University, Beijing, China
Chengchun Shi
Chengchun Shi
London School of Economics and Political Science
Large Language ModelsReinforcement LearningStatistics