occupancy-weighted value estimation

Design and implement policy-value estimators that reweight observed rewards or fitted Q-values by estimated state–action occupancy (importance) ratios to correct for distribution shift between the data-collection policy and a target policy. This includes building occupancy-weighted fitted Q-evaluation procedures, reward-reweighted estimators and doubly-robust variants, estimating the required density/occupancy ratios, and analyzing the resulting bias–variance tradeoffs in offline policy evaluation.

occupancy-weightedvalueestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.27
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the policy evaluation bias in offline reinforcement learning caused by distributional shift by proposing Fitted Occupancy Ratio Estimation (FORE). The method formulates a fixed-point equation for the occupancy ratio via an adjoint Bellman recursion and iteratively optimizes a density ratio objective under KL divergence using single-step transition data. Its key innovation lies in establishing the realizability of the occupancy ratio as a sufficient condition, thereby circumventing reliance on strong assumptions such as Bellman completeness. Theoretical analysis shows that FORE converges to the true occupancy ratio under KL divergence and achieves a finite-sample regret bound solely under this realizability condition. Moreover, the approach accommodates multiple value estimation schemes and exhibits both double robustness and strong empirical performance.

Bellman completenessdistribution shiftoccupancy ratio

Online Estimation and Inference for Robust Policy Evaluation in Reinforcement Learning

Oct 04, 2023
WL
Weidong Liu
🏛️ Shanghai Jiao Tong University | Shanghai University of Finance and Economics | Purdue University | New York University

In online policy evaluation for reinforcement learning, outliers and heavy-tailed reward distributions severely degrade the accuracy of parameter and value function estimates. Method: We propose the first unified robust online statistical inference framework, introducing Bahadur-type expansions to temporal difference (TD) learning—enabling incrementally updated, asymptotically normal estimators—and integrating robust statistical estimation with online variance adaptation. Contribution/Results: Theoretically, we establish asymptotic normality of the estimator under both heavy-tailed rewards and adversarial contamination. Empirically, our method significantly improves estimation stability and confidence interval coverage in both synthetic and real-world RL tasks; it achieves over 40% higher robustness against interference compared to standard TD methods, providing a new paradigm for robust policy evaluation that is both theoretically grounded and computationally feasible.

Addresses outlier contamination and heavy-tailed reward issues.Develops online robust policy evaluation in reinforcement learning.Provides statistical inference for model parameters and value functions.

Post Reinforcement Learning Inference

Feb 17, 2023
VS
Vasilis Syrgkanis
🏛️ Stanford University | Hong Kong University of Science and Technology

In reinforcement learning, adaptive interaction data—where the behavior policy is nonstationary—invalidates standard estimators, undermining asymptotic normality for off-policy counterfactual policy evaluation and dynamic treatment effect (DTE) inference. To address this, we propose a weighted Z-estimation framework that constructs time-varying adaptive weights to stabilize heteroskedasticity, achieving, for the first time in the RL off-policy setting, both consistent and asymptotically normal DTE estimation. Our approach integrates dynamic causal inference with asymptotic statistical theory, enabling rigorous hypothesis testing and construction of uniformly valid confidence regions. Simulation studies and real-world RL experiments demonstrate substantial improvements in confidence interval coverage and statistical power. The method provides the first solution for structural parameter inference under adaptive experimentation that simultaneously offers theoretical guarantees—namely consistency, asymptotic normality, and uniform validity—and empirical robustness.

Address nonstationary variance in adaptive reinforcement learning environmentsDevelop weighted Z-estimation for dynamic treatment effect analysisEstimate counterfactual policies post reinforcement learning data collection

Doubly Optimal Policy Evaluation for Reinforcement Learning

Oct 03, 2024
SL
Shuze Liu
🏛️ University of Virginia

High variance in policy evaluation for long-horizon reinforcement learning tasks—stemming from suboptimal behavior policies and baseline designs—severely limits sample efficiency. To address this, we propose a “doubly optimal” joint optimization framework that simultaneously optimizes both the behavior policy and the control variate baseline, achieving strict variance reduction while preserving unbiasedness. Our method integrates importance sampling, control variates, and policy gradients, and establishes a theoretical analysis grounded in the bias–variance trade-off. We prove that our estimator’s variance is strictly lower than that of all existing optimal estimators. Empirically, on multiple continuous-control benchmark tasks, our approach reduces estimation variance significantly under identical sample budgets, improves evaluation accuracy by over 40%, and achieves state-of-the-art (SOTA) performance.

Combines optimal data-collecting policy and data-processing baselineEnsures unbiased and lower variance than previous methodsReduces variance in reinforcement learning policy evaluation

Latest Papers

What's happening recently
View more

This work addresses off-policy evaluation in finite-horizon Markov decision processes under function approximation and limited data coverage. It proposes a novel method based on recursive reweighting and moment matching, which optimizes scalar weights through a value-function discriminator class in a top-down manner to align the reweighted returns with the expected return under the target policy. The approach unifies and generalizes existing techniques such as importance sampling and linear fitted Q-evaluation. Notably, under the sole assumption that the true Q-function is realizable within the chosen function class, it establishes the first finite-sample error bound that is independent of both the statistical complexity of the function class and the ambient dimensionality. This result advances the theoretical understanding of coverage conditions in offline reinforcement learning and significantly enhances the accuracy and robustness of policy evaluation.

finite-horizon MDPsmoment matchingoff-policy evaluation

This work addresses the challenge of Q-value overestimation in offline reinforcement learning, which arises from distributional shift between the behavior policy and the learning policy and can mislead policy optimization. To mitigate this issue, the paper proposes Behavior-Advantage Corrected Policy Evaluation (BAC-PE), a novel approach that leverages the Q-function of the behavior policy to correct the Q-estimates of the learning policy. BAC-PE further integrates a diffusion model for policy representation, enhancing expressiveness and training stability through Q-guided sampling and distribution-matching regularization. Theoretical analysis establishes the convergence of BAC-PE and provides an upper bound on its estimation error. Empirical evaluation on the D4RL benchmark demonstrates that the proposed method significantly outperforms existing offline reinforcement learning algorithms, achieving state-of-the-art performance.

behavioral policydistribution shiftoffline reinforcement learning

This work addresses the challenge of policy evaluation in offline reinforcement learning when immediate rewards are missing not at random (MNAR) due to sparse or truncated logging, which induces selection bias and undermines conventional evaluation methods. Focusing on finite-horizon Markov decision processes under MNAR reward missingness, the paper proposes a novel policy evaluation approach that avoids explicit modeling of the missingness mechanism. The method leverages future states as shadow variables to identify the conditional mean reward under complete data, integrating bridge functions within a Fitted-Q-Evaluation framework to construct a missingness-aware estimator for target policies. Theoretical analysis establishes the consistency of the proposed estimator and provides finite-sample error bounds. Empirical results demonstrate significant performance gains over existing methods in both simulated environments and real-world sepsis treatment data from the MIMIC-III database.

Markov Decision ProcessesMissing Not at RandomOff-Policy Evaluation

This work addresses the trade-off between off-policy bias and eligibility trace truncation in multi-step credit assignment within Q-learning. The authors propose an exploration-aware continuous gating mechanism that replaces conventional importance sampling, enabling adaptive interpolation between Watkins’s and Peng’s Q(λ) methods. This mechanism employs a state-action-dependent, differentiable gating function to dynamically modulate eligibility traces. Theoretically, the expected update operator under this approach is proven to be a contraction mapping with an exact fixed point, ensuring stability while allowing controlled balancing of bias and learning efficiency. Empirical results demonstrate that moderate gating effectively extends the credit assignment horizon and significantly accelerates early-stage learning, outperforming existing Q(λ) strategies that rely on extreme off-policy or on-policy assumptions.

credit assignmenteligibility tracesoff-policy bias

Hot Scholars

AW

Anqi Wu

Assistant Professor, Computational Science and Engineering, Georgia Tech
machine learningcomputational and statistical neuroscience
GG

Gennian Ge

Capital Normal University
CombinatoricsCoding theoryInformation Security
JM

Jianbiao Mei

Zhejiang University
computer visiondeep learning