Wasserstein Policy Learning for Distributional Outcomes

📅 2026-06-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses offline policy learning in settings where outcomes are probability distributions rather than scalars, aiming to optimize utility functionals defined via the Wasserstein barycenter. The study introduces Wasserstein geometry into this domain for the first time, establishing a policy learning framework tailored to distributional potential outcomes and integrating inverse probability weighting (IPW) with doubly robust (DR) estimators for policy optimization. By characterizing policy class complexity through Natarajan dimension and controlling uniform bias, the authors derive matching minimax lower bounds. Under the univariate Wasserstein setting, the resulting finite-sample regret bound achieves a leading term of Õ(√(Natarajan(Π)/N)), which is optimal in both sample size N and policy class complexity.
📝 Abstract
Offline policy learning has received growing attention in causal inference. The primary objective is to learn a policy (individualized treatment rule) as a mapping from covariates to treatment that maximizes the empirical welfare defined as the mean of scalar-valued potential outcomes. In this paper, we study offline policy learning with distribution-valued outcomes, where each potential outcome is a probability measure on $\mathbb{R}$ and the reward is defined through a utility functional applied to the Wasserstein barycenter of induced outcome distributions. We establish statistical guarantees for the policy learning framework based on both Inverse Probability Weighting (IPW) and Doubly Robust (DR) estimators. By handling the challenging uniform deviation over the product of the combinatorial policy class and the infinite-dimensional quantile domain, we prove that the finite-sample regret has leading dependence $\widetilde{\mathcal{O}}(\sqrt{\mathrm{N\text{-}dim}(Π)/N})$. In the one-dimensional Wasserstein setting and under the stated regularity conditions, the leading regret rate is still governed by the policy-class complexity. Moreover, we provide a minimax lower bound establishing the sharpness of the leading dependence on $N$ and $\mathrm{N\text{-}dim}(Π)$.
Problem

Research questions and friction points this paper is trying to address.

offline policy learning
distributional outcomes
Wasserstein barycenter
causal inference
individualized treatment rule
Innovation

Methods, ideas, or system contributions that make the work stand out.

distributional outcomes
Wasserstein barycenter
offline policy learning
doubly robust estimation
minimax regret
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yiyan Huang
School of Computing and Information Technology, Great Bay University, Guangdong, China
C
Cheuk Hang Leung
Department of Data Science, City University of Hong Kong, Hong Kong, China
Q
Qi Wu
Department of Data Science, City University of Hong Kong, Hong Kong, China
Z
Zhiheng Zhang
School of Statistics and Data Science & Institute of Big Data Research, Shanghai University of Finance and Economics, Shanghai, China