Score
Design and implement policy-value estimators that reweight observed rewards or fitted Q-values by estimated state–action occupancy (importance) ratios to correct for distribution shift between the data-collection policy and a target policy. This includes building occupancy-weighted fitted Q-evaluation procedures, reward-reweighted estimators and doubly-robust variants, estimating the required density/occupancy ratios, and analyzing the resulting bias–variance tradeoffs in offline policy evaluation.
This work addresses the policy evaluation bias in offline reinforcement learning caused by distributional shift by proposing Fitted Occupancy Ratio Estimation (FORE). The method formulates a fixed-point equation for the occupancy ratio via an adjoint Bellman recursion and iteratively optimizes a density ratio objective under KL divergence using single-step transition data. Its key innovation lies in establishing the realizability of the occupancy ratio as a sufficient condition, thereby circumventing reliance on strong assumptions such as Bellman completeness. Theoretical analysis shows that FORE converges to the true occupancy ratio under KL divergence and achieves a finite-sample regret bound solely under this realizability condition. Moreover, the approach accommodates multiple value estimation schemes and exhibits both double robustness and strong empirical performance.
In online policy evaluation for reinforcement learning, outliers and heavy-tailed reward distributions severely degrade the accuracy of parameter and value function estimates. Method: We propose the first unified robust online statistical inference framework, introducing Bahadur-type expansions to temporal difference (TD) learning—enabling incrementally updated, asymptotically normal estimators—and integrating robust statistical estimation with online variance adaptation. Contribution/Results: Theoretically, we establish asymptotic normality of the estimator under both heavy-tailed rewards and adversarial contamination. Empirically, our method significantly improves estimation stability and confidence interval coverage in both synthetic and real-world RL tasks; it achieves over 40% higher robustness against interference compared to standard TD methods, providing a new paradigm for robust policy evaluation that is both theoretically grounded and computationally feasible.
In reinforcement learning, adaptive interaction data—where the behavior policy is nonstationary—invalidates standard estimators, undermining asymptotic normality for off-policy counterfactual policy evaluation and dynamic treatment effect (DTE) inference. To address this, we propose a weighted Z-estimation framework that constructs time-varying adaptive weights to stabilize heteroskedasticity, achieving, for the first time in the RL off-policy setting, both consistent and asymptotically normal DTE estimation. Our approach integrates dynamic causal inference with asymptotic statistical theory, enabling rigorous hypothesis testing and construction of uniformly valid confidence regions. Simulation studies and real-world RL experiments demonstrate substantial improvements in confidence interval coverage and statistical power. The method provides the first solution for structural parameter inference under adaptive experimentation that simultaneously offers theoretical guarantees—namely consistency, asymptotic normality, and uniform validity—and empirical robustness.
High variance in policy evaluation for long-horizon reinforcement learning tasks—stemming from suboptimal behavior policies and baseline designs—severely limits sample efficiency. To address this, we propose a “doubly optimal” joint optimization framework that simultaneously optimizes both the behavior policy and the control variate baseline, achieving strict variance reduction while preserving unbiasedness. Our method integrates importance sampling, control variates, and policy gradients, and establishes a theoretical analysis grounded in the bias–variance trade-off. We prove that our estimator’s variance is strictly lower than that of all existing optimal estimators. Empirically, on multiple continuous-control benchmark tasks, our approach reduces estimation variance significantly under identical sample budgets, improves evaluation accuracy by over 40%, and achieves state-of-the-art (SOTA) performance.
This work addresses off-policy evaluation in finite-horizon Markov decision processes under function approximation and limited data coverage. It proposes a novel method based on recursive reweighting and moment matching, which optimizes scalar weights through a value-function discriminator class in a top-down manner to align the reweighted returns with the expected return under the target policy. The approach unifies and generalizes existing techniques such as importance sampling and linear fitted Q-evaluation. Notably, under the sole assumption that the true Q-function is realizable within the chosen function class, it establishes the first finite-sample error bound that is independent of both the statistical complexity of the function class and the ambient dimensionality. This result advances the theoretical understanding of coverage conditions in offline reinforcement learning and significantly enhances the accuracy and robustness of policy evaluation.
This work addresses the challenge of Q-value overestimation in offline reinforcement learning, which arises from distributional shift between the behavior policy and the learning policy and can mislead policy optimization. To mitigate this issue, the paper proposes Behavior-Advantage Corrected Policy Evaluation (BAC-PE), a novel approach that leverages the Q-function of the behavior policy to correct the Q-estimates of the learning policy. BAC-PE further integrates a diffusion model for policy representation, enhancing expressiveness and training stability through Q-guided sampling and distribution-matching regularization. Theoretical analysis establishes the convergence of BAC-PE and provides an upper bound on its estimation error. Empirical evaluation on the D4RL benchmark demonstrates that the proposed method significantly outperforms existing offline reinforcement learning algorithms, achieving state-of-the-art performance.
This work addresses the challenge of policy evaluation in offline reinforcement learning when immediate rewards are missing not at random (MNAR) due to sparse or truncated logging, which induces selection bias and undermines conventional evaluation methods. Focusing on finite-horizon Markov decision processes under MNAR reward missingness, the paper proposes a novel policy evaluation approach that avoids explicit modeling of the missingness mechanism. The method leverages future states as shadow variables to identify the conditional mean reward under complete data, integrating bridge functions within a Fitted-Q-Evaluation framework to construct a missingness-aware estimator for target policies. Theoretical analysis establishes the consistency of the proposed estimator and provides finite-sample error bounds. Empirical results demonstrate significant performance gains over existing methods in both simulated environments and real-world sepsis treatment data from the MIMIC-III database.
This work addresses the trade-off between off-policy bias and eligibility trace truncation in multi-step credit assignment within Q-learning. The authors propose an exploration-aware continuous gating mechanism that replaces conventional importance sampling, enabling adaptive interpolation between Watkins’s and Peng’s Q(λ) methods. This mechanism employs a state-action-dependent, differentiable gating function to dynamically modulate eligibility traces. Theoretically, the expected update operator under this approach is proven to be a contraction mapping with an exact fixed point, ensuring stability while allowing controlled balancing of bias and learning efficiency. Empirical results demonstrate that moderate gating effectively extends the credit assignment horizon and significantly accelerates early-stage learning, outperforming existing Q(λ) strategies that rely on extreme off-policy or on-policy assumptions.