Score
Designs, builds, or analyzes estimators that use logged data gathered under one behavior policy to estimate the value or average outcome of a different target policy by reweighting observed outcomes with propensity/exposure ratios or by modeling value functions (e.g., fitted Q‑evaluation). This work covers importance‑sampling and kernelized IPS estimators, corrections for missingness including MNAR and recovered rewards, and the derivation of consistency and finite‑sample error bounds for those estimators.
This paper addresses robust policy evaluation and learning under distributional shift using offline observational data with continuous interventions, relaxing the conventional assumptions of discrete interventions and covariate shift invariance. Methodologically, it introduces distributionally robust optimization (DRO) into the continuous-intervention setting for the first time; proposes a kernel-based inverse probability weighting (IPW) estimator to mitigate sample sparsity and exclusion bias arising from continuous actions; and establishes finite-sample convergence guarantees. Theoretical analysis proves the consistency of the proposed estimator. Empirical results demonstrate that the method significantly outperforms standard IPW and existing baselines across diverse distribution shifts, yielding improvements in both estimation accuracy and policy learning stability.
This study addresses the challenge of learning treatment assignment policies under missing data, proposing a robust framework that mitigates bias and suboptimal decisions arising from ignoring the missingness mechanism. The authors extend efficient estimation of the average treatment effect (ATE) to settings with missing-at-random (MAR) and missing completely at random conditional on covariates and treatment (MCCAR) assumptions. Leveraging semiparametric efficiency theory, they develop a doubly robust estimator that jointly models the missingness mechanism and the conditional average treatment effect (CATE), thereby effectively utilizing partially observed samples. A key contribution is the first systematic comparison of estimation efficiency for policy value under MAR versus MCCAR, with theoretical proof that the MAR-based estimator remains more efficient even when MCCAR holds. Experiments demonstrate that the proposed method achieves near-oracle performance when the missingness mechanism is correctly specified, whereas misspecified approaches incur substantial bias.
This paper investigates a counterintuitive phenomenon in off-policy evaluation (OPE): even when the true behavior policy is Markovian, employing history-dependent (non-Markovian) estimators in importance sampling can reduce mean squared error (MSE). We derive the first rigorous bias–variance decomposition for importance sampling estimators in OPE and prove that history-dependent modeling substantially reduces asymptotic variance—monotonically decreasing with increasing history length. Methodologically, we integrate sequential importance sampling, marginalized weights, doubly robust estimation, and both parametric and nonparametric behavior policy modeling. Theoretically and empirically, we demonstrate that history-dependent estimators effectively balance bias and variance in finite-sample regimes, yielding significant gains in OPE accuracy. Our work establishes a novel paradigm for behavior policy modeling in OPE, challenging the conventional assumption that Markovian approximations are always optimal.
In reinforcement learning, adaptive interaction data—where the behavior policy is nonstationary—invalidates standard estimators, undermining asymptotic normality for off-policy counterfactual policy evaluation and dynamic treatment effect (DTE) inference. To address this, we propose a weighted Z-estimation framework that constructs time-varying adaptive weights to stabilize heteroskedasticity, achieving, for the first time in the RL off-policy setting, both consistent and asymptotically normal DTE estimation. Our approach integrates dynamic causal inference with asymptotic statistical theory, enabling rigorous hypothesis testing and construction of uniformly valid confidence regions. Simulation studies and real-world RL experiments demonstrate substantial improvements in confidence interval coverage and statistical power. The method provides the first solution for structural parameter inference under adaptive experimentation that simultaneously offers theoretical guarantees—namely consistency, asymptotic normality, and uniform validity—and empirical robustness.
Learning continuous treatment policies from observational data faces three key challenges: nonparametric welfare estimation, infinite-dimensional policy spaces, and shape constraints (e.g., monotonicity or convexity). Method: We propose a novel paradigm that approximates the shape-constrained policy space via a sequence of finite-dimensional subspaces; we develop a data-adaptive tuned penalization algorithm, integrating kernel-based welfare estimation, machine learning–based propensity score modeling, regularized optimization, and constrained function approximation. Contribution/Results: We establish, for the first time, an oracle inequality for welfare regret under continuous treatments. Theoretically, our estimator achieves statistically optimal convergence rates under both known and unknown propensity scores. Empirically, it significantly enhances out-of-sample policy extrapolation robustness and real-world effectiveness.
This study addresses the limitations of existing mean-difference estimators in randomized controlled trials, which, while unbiased, often exhibit limited precision and lack a systematic framework for selection tailored to distinct analytical objectives—such as statistical inference versus decision-making. To overcome this, the authors propose a general sample-splitting framework that reframes estimator selection as a goal-oriented comparative task, thereby avoiding the validity risks associated with post hoc choices. By estimating the distributions of metrics like mean squared error and regret across multiple estimation methods—including covariate adjustment, weighted least squares, and simple difference-in-means—the framework enables objective evaluation. Empirical analyses on Amazon supply-chain experiments and the Strengthening Democracy Challenge (25 interventions) reveal that weighted least squares performs best for inference, whereas difference-in-means yields the lowest regret in decision contexts, demonstrating for the first time that the optimal estimator varies systematically with the analytical goal.
This work addresses the challenge of policy evaluation in offline reinforcement learning when immediate rewards are missing not at random (MNAR) due to sparse or truncated logging, which induces selection bias and undermines conventional evaluation methods. Focusing on finite-horizon Markov decision processes under MNAR reward missingness, the paper proposes a novel policy evaluation approach that avoids explicit modeling of the missingness mechanism. The method leverages future states as shadow variables to identify the conditional mean reward under complete data, integrating bridge functions within a Fitted-Q-Evaluation framework to construct a missingness-aware estimator for target policies. Theoretical analysis establishes the consistency of the proposed estimator and provides finite-sample error bounds. Empirical results demonstrate significant performance gains over existing methods in both simulated environments and real-world sepsis treatment data from the MIMIC-III database.
This work addresses the challenge of off-policy evaluation under right-censored survival outcomes, where existing methods are prone to systematic bias and yield inaccurate policy value estimates. To mitigate this issue, the study introduces inverse probability of censoring weighting (IPCW) into off-policy evaluation for the first time, proposing two novel estimators—IPCW-IPS and IPCW-DR—that are both unbiased and doubly robust, effectively correcting for censoring-induced bias. Furthermore, the proposed framework naturally extends to policy optimization under budget constraints. Experimental results on both synthetic and real-world datasets demonstrate that the method substantially improves the accuracy of policy evaluation and enhances learning performance in censored environments.
This work addresses the challenge of statistical inference in contextual bandits when the outcome model is misspecified, a setting in which existing algorithms like LinUCB can yield non-Gaussian estimators and invalid confidence intervals. The authors propose a Z-estimation framework based on inverse probability weighting that accommodates a broad class of marginal moment objectives. They introduce a novel stability condition—scaled inverse propensity convergence—and establish, for the first time, sufficient conditions guaranteeing consistency and asymptotic normality for a wide range of adaptive policies. By integrating sandwich variance estimation with policy stability theory, the method supports both multi-armed and smooth contextual allocation strategies. Empirical evaluations on synthetic data and the HeartSteps V1 trial demonstrate accurate confidence interval coverage and competitive estimation performance across multiple objectives.