doubly robust policy learning

Designs and implements estimators and learning algorithms for decision policies that combine outcome models and propensity (inverse‑probability) models to produce doubly robust policy value estimates and policy learners. Builds estimators for policy evaluation including distributional outcomes and quantiles, and analyzes them by proving uniform deviation bounds over policy classes and deriving regret guarantees that are robust to certain forms of model misspecification.

doublyrobustpolicylearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Online Estimation and Inference for Robust Policy Evaluation in Reinforcement Learning

Oct 04, 2023
WL
Weidong Liu
🏛️ Shanghai Jiao Tong University | Shanghai University of Finance and Economics | Purdue University | New York University

In online policy evaluation for reinforcement learning, outliers and heavy-tailed reward distributions severely degrade the accuracy of parameter and value function estimates. Method: We propose the first unified robust online statistical inference framework, introducing Bahadur-type expansions to temporal difference (TD) learning—enabling incrementally updated, asymptotically normal estimators—and integrating robust statistical estimation with online variance adaptation. Contribution/Results: Theoretically, we establish asymptotic normality of the estimator under both heavy-tailed rewards and adversarial contamination. Empirically, our method significantly improves estimation stability and confidence interval coverage in both synthetic and real-world RL tasks; it achieves over 40% higher robustness against interference compared to standard TD methods, providing a new paradigm for robust policy evaluation that is both theoretically grounded and computationally feasible.

Addresses outlier contamination and heavy-tailed reward issues.Develops online robust policy evaluation in reinforcement learning.Provides statistical inference for model parameters and value functions.

Policy Learning with Distributional Welfare

Nov 27, 2023
YC
Yifan Cui
🏛️ Zhejiang University | University of Bristol

This paper addresses the vulnerability of conventional utilitarian policy learning—based on the conditional average treatment effect (CATE)—to outliers and its inability to flexibly accommodate policy caution or leniency under heterogeneous individual treatment effects. We propose an optimal intervention allocation framework grounded in the conditional quantile treatment effect (QoTE). To our knowledge, this is the first work to incorporate distributionally robust welfare into policy learning, formulating a minimax strategy based on QoTE that accommodates the fundamental challenge of non-point-identification of the counterfactual joint distribution. Our method integrates causal inference, distributionally robust optimization, and decision theory, supporting both stochastic and deterministic policies under various identification assumptions. We establish an asymptotically tight upper bound on the regret and demonstrate robustness to model misspecification. The framework generalizes to any welfare objective defined as a functional of the potential outcomes’ joint distribution.

Generalizing welfare frameworks beyond utilitarian criteria for heterogeneous populationsIdentifying policies robust to uncertainty in counterfactual outcome distributionsOptimal treatment allocation targeting distributional welfare, not just average effects

Policy Learning with New Treatments

Oct 10, 2022
SH
Samuel Higbee
🏛️ University of Chicago

This paper addresses optimal policy assignment under partial treatment coverage in heterogeneous populations, where experiments only implement a subset of possible treatment values, limiting generalizability to untested interventions. Method: We propose the first framework integrating shape-constrained partial identification of treatment effects—incorporating monotonicity or convexity constraints—with a minimax regret criterion, formulated as a tractable mixed-integer linear program (MILP). The method combines nonparametric conditional average treatment effect (CATE) estimation, shape restrictions, minimax optimization, and efficient linear/integer programming solvers. Contribution/Results: Applied to a Kenyan rural electricity subsidy experiment, our framework recommends novel, experimentally untested treatment levels—covering nearly the entire population—while reducing maximum regret by over 60%. It substantially enhances policy extrapolation capability and robustness beyond conventional methods constrained to observed treatment supports.

Allocating treatments not tested in experiments to heterogeneous populationsMinimizing maximum regret in policy decisions using linear programmingPartially identifying new treatment effects via shape restrictions

This paper addresses off-policy evaluation (OPE) in Markov decision processes (MDPs) under weak distributional overlap. Conventional OPE methods rely on strong overlap assumptions—namely, boundedness of the stationary density ratio between target and behavior policies—which often fail in unbounded state spaces. We propose the Truncated Doubly Robust (TDR) estimator, the first to achieve statistical consistency under the significantly weaker condition that the density ratio is merely square-integrable. Under this condition, TDR attains a $sqrt{T}$-consistency rate; otherwise, it achieves the minimax-optimal slower rate. Our theoretical analysis rigorously establishes consistency and optimality by jointly characterizing stationary distributions and modeling mixing properties. Empirical results demonstrate that TDR substantially improves estimation accuracy under weak overlap, outperforming existing OPE methods.

Achieving minimax convergence rates without bounded state spacesDeveloping truncated doubly robust estimators for off-policy evaluationEvaluating policies in MDPs with weak distributional overlap

Doubly Optimal Policy Evaluation for Reinforcement Learning

Oct 03, 2024
SL
Shuze Liu
🏛️ University of Virginia

High variance in policy evaluation for long-horizon reinforcement learning tasks—stemming from suboptimal behavior policies and baseline designs—severely limits sample efficiency. To address this, we propose a “doubly optimal” joint optimization framework that simultaneously optimizes both the behavior policy and the control variate baseline, achieving strict variance reduction while preserving unbiasedness. Our method integrates importance sampling, control variates, and policy gradients, and establishes a theoretical analysis grounded in the bias–variance trade-off. We prove that our estimator’s variance is strictly lower than that of all existing optimal estimators. Empirically, on multiple continuous-control benchmark tasks, our approach reduces estimation variance significantly under identical sample budgets, improves evaluation accuracy by over 40%, and achieves state-of-the-art (SOTA) performance.

Combines optimal data-collecting policy and data-processing baselineEnsures unbiased and lower variance than previous methodsReduces variance in reinforcement learning policy evaluation

Latest Papers

What's happening recently
View more

This work addresses offline policy learning in settings where outcomes are probability distributions rather than scalars, aiming to optimize utility functionals defined via the Wasserstein barycenter. The study introduces Wasserstein geometry into this domain for the first time, establishing a policy learning framework tailored to distributional potential outcomes and integrating inverse probability weighting (IPW) with doubly robust (DR) estimators for policy optimization. By characterizing policy class complexity through Natarajan dimension and controlling uniform bias, the authors derive matching minimax lower bounds. Under the univariate Wasserstein setting, the resulting finite-sample regret bound achieves a leading term of Õ(√(Natarajan(Π)/N)), which is optimal in both sample size N and policy class complexity.

causal inferencedistributional outcomesindividualized treatment rule

Semiparametric Efficiency in Policy Learning with General Treatments

Dec 22, 2025
YF
Yue Fang
🏛️ The Chinese University of Hong Kong, Shenzhen | University of Southern California | Peking University

This paper addresses policy learning with general treatment variables—discrete, continuous, or mixed—and establishes the first unified semiparametrically efficient estimation framework. It systematically characterizes the asymptotic efficiency bounds for welfare regret under both deterministic and stochastic policies. Methodologically, it introduces a novel efficiency definition for welfare regret, overcoming the non-differentiability barrier inherent in conventional parametric paths for deterministic policies; identifies a new “Higher-Order Efficiency” (HIR) phenomenon wherein inverse-probability-weighted (IPW) estimators using estimated rather than true propensity scores achieve superior asymptotic efficiency; and derives the asymptotic distributions of mainstream policy estimators via pathwise differentiability analysis, convolution theorem, and IPW. Empirically, the framework significantly improves estimation efficiency in job training and savings program evaluations and reveals mean-shift effects.

Analyzes asymptotic distribution of welfare regret and efficiency of policy estimatorsCharacterizes semiparametric efficiency for policy learning with general treatmentsEstablishes efficiency bounds for randomized policies under known and estimated propensities

When data are insufficient to learn policies with low regret or significantly better performance than a baseline, how can we characterize the intrinsic difficulty of policy learning? This work proposes a unified framework to systematically study three fundamental problems: optimal policy learning, improved policy learning, and policy existence verification. Through theoretical analysis, problem reductions, and sample complexity comparisons, the paper establishes a strict or partially strict hierarchy of difficulty among these tasks: optimal policy learning is provably harder than improved policy learning, and under natural conditions, a sublinear polynomial complexity gap separates improved policy learning from existence verification. Notably, this study formalizes the policy existence problem for the first time and reveals that even when constructing an improved policy is infeasible, efficiently determining its existence may still be possible.

improving policyoptimal policypolicy existence

This work addresses the high variance in reweighted estimators arising from excessively small propensity scores in offline policy learning. To mitigate this issue, the authors propose a novel weight-clipping algorithm that adaptively selects the clipping threshold by minimizing the mean squared error of the policy value estimator. The resulting bilevel discontinuous optimization problem is reformulated as a Heaviside-composite optimization problem and efficiently solved via asymptotic integer programming. Theoretical analysis provides an upper bound on policy suboptimality, ensuring robust performance guarantees. Empirical results demonstrate that the proposed method substantially reduces estimation variance and improves the practical performance of learned policies. Overall, this study offers a rigorous and computationally tractable optimization framework for offline policy learning with discontinuous objectives.

high varianceoffline policy learningpolicy value estimation

This study addresses the limitation of traditional conditional treatment effect estimation, which is confined to scalar outcomes and struggles with multidimensional, ordinal, or merely rankable preference-based outcomes. The authors propose a Preference-based Conditional Treatment Effect (CPTE) framework that characterizes heterogeneous treatment effects by modeling the ordinal relationships among outcomes and enables policy learning. This framework unifies various preference-driven causal estimands for the first time, establishes novel identifiability conditions with clear interpretability, and integrates matching, quantile regression, and distributional regression to construct an efficient influence-function-based estimator. This estimator corrects plug-in bias and enhances policy value. Experiments on synthetic and semi-synthetic data demonstrate that the proposed method significantly outperforms existing approaches, confirming its effectiveness and practical potential.

Conditional Treatment EffectHeterogeneous Treatment EffectsNon-identifiability

Hot Scholars

VK

Vinod K. Sangwan

Northwestern University; University of Maryland College Park; Indian Institute of Technology Bombay
Condensed matter physicsNeuromorphic computingQuantum materialsEnergy conversion and storage
SZ

Shuxin Zheng

Deputy Director, Zhongguancun Institute of Artificial Intelligence
General AIGenerative AI
YS

Yucheng Shi

University of Georgia
Synthetic DataData-centric AIResponsible AIExplainability
JW

Jannis Weil

Leibniz University Hannover
Reinforcement LearningNetworksCommunication SystemsMultimedia Systems