Score
Designs and implements estimators and learning algorithms for decision policies that combine outcome models and propensity (inverse‑probability) models to produce doubly robust policy value estimates and policy learners. Builds estimators for policy evaluation including distributional outcomes and quantiles, and analyzes them by proving uniform deviation bounds over policy classes and deriving regret guarantees that are robust to certain forms of model misspecification.
In online policy evaluation for reinforcement learning, outliers and heavy-tailed reward distributions severely degrade the accuracy of parameter and value function estimates. Method: We propose the first unified robust online statistical inference framework, introducing Bahadur-type expansions to temporal difference (TD) learning—enabling incrementally updated, asymptotically normal estimators—and integrating robust statistical estimation with online variance adaptation. Contribution/Results: Theoretically, we establish asymptotic normality of the estimator under both heavy-tailed rewards and adversarial contamination. Empirically, our method significantly improves estimation stability and confidence interval coverage in both synthetic and real-world RL tasks; it achieves over 40% higher robustness against interference compared to standard TD methods, providing a new paradigm for robust policy evaluation that is both theoretically grounded and computationally feasible.
This paper addresses the vulnerability of conventional utilitarian policy learning—based on the conditional average treatment effect (CATE)—to outliers and its inability to flexibly accommodate policy caution or leniency under heterogeneous individual treatment effects. We propose an optimal intervention allocation framework grounded in the conditional quantile treatment effect (QoTE). To our knowledge, this is the first work to incorporate distributionally robust welfare into policy learning, formulating a minimax strategy based on QoTE that accommodates the fundamental challenge of non-point-identification of the counterfactual joint distribution. Our method integrates causal inference, distributionally robust optimization, and decision theory, supporting both stochastic and deterministic policies under various identification assumptions. We establish an asymptotically tight upper bound on the regret and demonstrate robustness to model misspecification. The framework generalizes to any welfare objective defined as a functional of the potential outcomes’ joint distribution.
This paper addresses optimal policy assignment under partial treatment coverage in heterogeneous populations, where experiments only implement a subset of possible treatment values, limiting generalizability to untested interventions. Method: We propose the first framework integrating shape-constrained partial identification of treatment effects—incorporating monotonicity or convexity constraints—with a minimax regret criterion, formulated as a tractable mixed-integer linear program (MILP). The method combines nonparametric conditional average treatment effect (CATE) estimation, shape restrictions, minimax optimization, and efficient linear/integer programming solvers. Contribution/Results: Applied to a Kenyan rural electricity subsidy experiment, our framework recommends novel, experimentally untested treatment levels—covering nearly the entire population—while reducing maximum regret by over 60%. It substantially enhances policy extrapolation capability and robustness beyond conventional methods constrained to observed treatment supports.
This paper addresses off-policy evaluation (OPE) in Markov decision processes (MDPs) under weak distributional overlap. Conventional OPE methods rely on strong overlap assumptions—namely, boundedness of the stationary density ratio between target and behavior policies—which often fail in unbounded state spaces. We propose the Truncated Doubly Robust (TDR) estimator, the first to achieve statistical consistency under the significantly weaker condition that the density ratio is merely square-integrable. Under this condition, TDR attains a $sqrt{T}$-consistency rate; otherwise, it achieves the minimax-optimal slower rate. Our theoretical analysis rigorously establishes consistency and optimality by jointly characterizing stationary distributions and modeling mixing properties. Empirical results demonstrate that TDR substantially improves estimation accuracy under weak overlap, outperforming existing OPE methods.
High variance in policy evaluation for long-horizon reinforcement learning tasks—stemming from suboptimal behavior policies and baseline designs—severely limits sample efficiency. To address this, we propose a “doubly optimal” joint optimization framework that simultaneously optimizes both the behavior policy and the control variate baseline, achieving strict variance reduction while preserving unbiasedness. Our method integrates importance sampling, control variates, and policy gradients, and establishes a theoretical analysis grounded in the bias–variance trade-off. We prove that our estimator’s variance is strictly lower than that of all existing optimal estimators. Empirically, on multiple continuous-control benchmark tasks, our approach reduces estimation variance significantly under identical sample budgets, improves evaluation accuracy by over 40%, and achieves state-of-the-art (SOTA) performance.
This work addresses offline policy learning in settings where outcomes are probability distributions rather than scalars, aiming to optimize utility functionals defined via the Wasserstein barycenter. The study introduces Wasserstein geometry into this domain for the first time, establishing a policy learning framework tailored to distributional potential outcomes and integrating inverse probability weighting (IPW) with doubly robust (DR) estimators for policy optimization. By characterizing policy class complexity through Natarajan dimension and controlling uniform bias, the authors derive matching minimax lower bounds. Under the univariate Wasserstein setting, the resulting finite-sample regret bound achieves a leading term of Õ(√(Natarajan(Π)/N)), which is optimal in both sample size N and policy class complexity.
This paper addresses policy learning with general treatment variables—discrete, continuous, or mixed—and establishes the first unified semiparametrically efficient estimation framework. It systematically characterizes the asymptotic efficiency bounds for welfare regret under both deterministic and stochastic policies. Methodologically, it introduces a novel efficiency definition for welfare regret, overcoming the non-differentiability barrier inherent in conventional parametric paths for deterministic policies; identifies a new “Higher-Order Efficiency” (HIR) phenomenon wherein inverse-probability-weighted (IPW) estimators using estimated rather than true propensity scores achieve superior asymptotic efficiency; and derives the asymptotic distributions of mainstream policy estimators via pathwise differentiability analysis, convolution theorem, and IPW. Empirically, the framework significantly improves estimation efficiency in job training and savings program evaluations and reveals mean-shift effects.
When data are insufficient to learn policies with low regret or significantly better performance than a baseline, how can we characterize the intrinsic difficulty of policy learning? This work proposes a unified framework to systematically study three fundamental problems: optimal policy learning, improved policy learning, and policy existence verification. Through theoretical analysis, problem reductions, and sample complexity comparisons, the paper establishes a strict or partially strict hierarchy of difficulty among these tasks: optimal policy learning is provably harder than improved policy learning, and under natural conditions, a sublinear polynomial complexity gap separates improved policy learning from existence verification. Notably, this study formalizes the policy existence problem for the first time and reveals that even when constructing an improved policy is infeasible, efficiently determining its existence may still be possible.
This work addresses the high variance in reweighted estimators arising from excessively small propensity scores in offline policy learning. To mitigate this issue, the authors propose a novel weight-clipping algorithm that adaptively selects the clipping threshold by minimizing the mean squared error of the policy value estimator. The resulting bilevel discontinuous optimization problem is reformulated as a Heaviside-composite optimization problem and efficiently solved via asymptotic integer programming. Theoretical analysis provides an upper bound on policy suboptimality, ensuring robust performance guarantees. Empirical results demonstrate that the proposed method substantially reduces estimation variance and improves the practical performance of learned policies. Overall, this study offers a rigorous and computationally tractable optimization framework for offline policy learning with discontinuous objectives.
This study addresses the limitation of traditional conditional treatment effect estimation, which is confined to scalar outcomes and struggles with multidimensional, ordinal, or merely rankable preference-based outcomes. The authors propose a Preference-based Conditional Treatment Effect (CPTE) framework that characterizes heterogeneous treatment effects by modeling the ordinal relationships among outcomes and enables policy learning. This framework unifies various preference-driven causal estimands for the first time, establishes novel identifiability conditions with clear interpretability, and integrates matching, quantile regression, and distributional regression to construct an efficient influence-function-based estimator. This estimator corrects plug-in bias and enhances policy value. Experiments on synthetic and semi-synthetic data demonstrate that the proposed method significantly outperforms existing approaches, confirming its effectiveness and practical potential.