off-policy estimation

Designs, builds, or analyzes estimators that use logged data gathered under one behavior policy to estimate the value or average outcome of a different target policy by reweighting observed outcomes with propensity/exposure ratios or by modeling value functions (e.g., fitted Q‑evaluation). This work covers importance‑sampling and kernelized IPS estimators, corrections for missingness including MNAR and recovered rewards, and the derivation of consistency and finite‑sample error bounds for those estimators.

off-policyestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Distributionally Robust Policy Evaluation and Learning for Continuous Treatment with Observational Data

Jan 18, 2025
CH
Cheuk Hang Leung
🏛️ City University of Hong Kong | The Hong Kong Polytechnic University

This paper addresses robust policy evaluation and learning under distributional shift using offline observational data with continuous interventions, relaxing the conventional assumptions of discrete interventions and covariate shift invariance. Methodologically, it introduces distributionally robust optimization (DRO) into the continuous-intervention setting for the first time; proposes a kernel-based inverse probability weighting (IPW) estimator to mitigate sample sparsity and exclusion bias arising from continuous actions; and establishes finite-sample convergence guarantees. Theoretical analysis proves the consistency of the proposed estimator. Empirical results demonstrate that the method significantly outperforms standard IPW and existing baselines across diverse distribution shifts, yielding improvements in both estimation accuracy and policy learning stability.

Continuous ProcessingData Distribution ChangesPolicy Learning

This study addresses the challenge of learning treatment assignment policies under missing data, proposing a robust framework that mitigates bias and suboptimal decisions arising from ignoring the missingness mechanism. The authors extend efficient estimation of the average treatment effect (ATE) to settings with missing-at-random (MAR) and missing completely at random conditional on covariates and treatment (MCCAR) assumptions. Leveraging semiparametric efficiency theory, they develop a doubly robust estimator that jointly models the missingness mechanism and the conditional average treatment effect (CATE), thereby effectively utilizing partially observed samples. A key contribution is the first systematic comparison of estimation efficiency for policy value under MAR versus MCCAR, with theoretical proof that the MAR-based estimator remains more efficient even when MCCAR holds. Experiments demonstrate that the proposed method achieves near-oracle performance when the missingness mechanism is correctly specified, whereas misspecified approaches incur substantial bias.

conditional average treatment effectmissing at randommissing treatment data

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation

May 28, 2025
HZ
Hongyi Zhou
🏛️ Tsinghua University | University of Wisconsin – Madison | London School of Economics and Political Science

This paper investigates a counterintuitive phenomenon in off-policy evaluation (OPE): even when the true behavior policy is Markovian, employing history-dependent (non-Markovian) estimators in importance sampling can reduce mean squared error (MSE). We derive the first rigorous bias–variance decomposition for importance sampling estimators in OPE and prove that history-dependent modeling substantially reduces asymptotic variance—monotonically decreasing with increasing history length. Methodologically, we integrate sequential importance sampling, marginalized weights, doubly robust estimation, and both parametric and nonparametric behavior policy modeling. Theoretically and empirically, we demonstrate that history-dependent estimators effectively balance bias and variance in finite-sample regimes, yielding significant gains in OPE accuracy. Our work establishes a novel paradigm for behavior policy modeling in OPE, challenging the conventional assumption that Markovian approximations are always optimal.

Analyzes bias-variance trade-off in importance sampling estimatorsExplains why history-dependent behavior policy reduces MSE in OPEExtends findings to various OPE estimators and estimation methods

Post Reinforcement Learning Inference

Feb 17, 2023
VS
Vasilis Syrgkanis
🏛️ Stanford University | Hong Kong University of Science and Technology

In reinforcement learning, adaptive interaction data—where the behavior policy is nonstationary—invalidates standard estimators, undermining asymptotic normality for off-policy counterfactual policy evaluation and dynamic treatment effect (DTE) inference. To address this, we propose a weighted Z-estimation framework that constructs time-varying adaptive weights to stabilize heteroskedasticity, achieving, for the first time in the RL off-policy setting, both consistent and asymptotically normal DTE estimation. Our approach integrates dynamic causal inference with asymptotic statistical theory, enabling rigorous hypothesis testing and construction of uniformly valid confidence regions. Simulation studies and real-world RL experiments demonstrate substantial improvements in confidence interval coverage and statistical power. The method provides the first solution for structural parameter inference under adaptive experimentation that simultaneously offers theoretical guarantees—namely consistency, asymptotic normality, and uniform validity—and empirical robustness.

Address nonstationary variance in adaptive reinforcement learning environmentsDevelop weighted Z-estimation for dynamic treatment effect analysisEstimate counterfactual policies post reinforcement learning data collection

Data-driven Policy Learning for Continuous Treatments

Feb 04, 2024
CA
Chunrong Ai
🏛️ The Chinese University of Hong Kong, Shenzhen | Peking University

Learning continuous treatment policies from observational data faces three key challenges: nonparametric welfare estimation, infinite-dimensional policy spaces, and shape constraints (e.g., monotonicity or convexity). Method: We propose a novel paradigm that approximates the shape-constrained policy space via a sequence of finite-dimensional subspaces; we develop a data-adaptive tuned penalization algorithm, integrating kernel-based welfare estimation, machine learning–based propensity score modeling, regularized optimization, and constrained function approximation. Contribution/Results: We establish, for the first time, an oracle inequality for welfare regret under continuous treatments. Theoretically, our estimator achieves statistically optimal convergence rates under both known and unknown propensity scores. Empirically, it significantly enhances out-of-sample policy extrapolation robustness and real-world effectiveness.

Addressing infinite-dimensional policy spaces with shape restrictionsAutomating tuning parameters for empirical welfare maximizationLearning optimal continuous treatment policies from observational data

Latest Papers

What's happening recently
View more

This study addresses the limitations of existing mean-difference estimators in randomized controlled trials, which, while unbiased, often exhibit limited precision and lack a systematic framework for selection tailored to distinct analytical objectives—such as statistical inference versus decision-making. To overcome this, the authors propose a general sample-splitting framework that reframes estimator selection as a goal-oriented comparative task, thereby avoiding the validity risks associated with post hoc choices. By estimating the distributions of metrics like mean squared error and regret across multiple estimation methods—including covariate adjustment, weighted least squares, and simple difference-in-means—the framework enables objective evaluation. Empirical analyses on Amazon supply-chain experiments and the Strengthening Democracy Challenge (25 interventions) reveal that weighted least squares performs best for inference, whereas difference-in-means yields the lowest regret in decision contexts, demonstrating for the first time that the optimal estimator varies systematically with the analytical goal.

estimator selectionheavy-tailed outcomesheterogeneous treatment effects

This work addresses the challenge of policy evaluation in offline reinforcement learning when immediate rewards are missing not at random (MNAR) due to sparse or truncated logging, which induces selection bias and undermines conventional evaluation methods. Focusing on finite-horizon Markov decision processes under MNAR reward missingness, the paper proposes a novel policy evaluation approach that avoids explicit modeling of the missingness mechanism. The method leverages future states as shadow variables to identify the conditional mean reward under complete data, integrating bridge functions within a Fitted-Q-Evaluation framework to construct a missingness-aware estimator for target policies. Theoretical analysis establishes the consistency of the proposed estimator and provides finite-sample error bounds. Empirical results demonstrate significant performance gains over existing methods in both simulated environments and real-world sepsis treatment data from the MIMIC-III database.

Markov Decision ProcessesMissing Not at RandomOff-Policy Evaluation

This work addresses the challenge of off-policy evaluation under right-censored survival outcomes, where existing methods are prone to systematic bias and yield inaccurate policy value estimates. To mitigate this issue, the study introduces inverse probability of censoring weighting (IPCW) into off-policy evaluation for the first time, proposing two novel estimators—IPCW-IPS and IPCW-DR—that are both unbiased and doubly robust, effectively correcting for censoring-induced bias. Furthermore, the proposed framework naturally extends to policy optimization under budget constraints. Experimental results on both synthetic and real-world datasets demonstrate that the method substantially improves the accuracy of policy evaluation and enhances learning performance in censored environments.

CensoringOff-Policy EvaluationPolicy Learning

This work addresses the challenge of statistical inference in contextual bandits when the outcome model is misspecified, a setting in which existing algorithms like LinUCB can yield non-Gaussian estimators and invalid confidence intervals. The authors propose a Z-estimation framework based on inverse probability weighting that accommodates a broad class of marginal moment objectives. They introduce a novel stability condition—scaled inverse propensity convergence—and establish, for the first time, sufficient conditions guaranteeing consistency and asymptotic normality for a wide range of adaptive policies. By integrating sandwich variance estimation with policy stability theory, the method supports both multi-armed and smooth contextual allocation strategies. Empirical evaluations on synthetic data and the HeartSteps V1 trial demonstrate accurate confidence interval coverage and competitive estimation performance across multiple objectives.

adaptive experimentscontextual banditsestimator stability

Hot Scholars

SL

Sergey Levine

UC Berkeley, Physical Intelligence
Machine LearningRoboticsReinforcement Learning
JH

Jianye Hao

Huawei Noah's Ark Lab/Tianjin University
Multiagent SystemsEmbodied AI
JZ

Jingren Zhou

Alibaba Group, Microsoft
Cloud ComputingLarge Scale Distributed SystemsMachine LearningQuery Processing
QL

Qiyang Li

University of California, Berkeley
Artificial Intelligence in RoboticsMachine Learning
DL

Dahua Lin

The Chinese University of Hong Kong
computer visionmachine learningprobabilistic inferencebayesian nonparametrics