Regularizing Discrimination in Optimal Policy Learning with Distributional Targets

๐Ÿ“… 2024-01-31
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF

career value

197K/year
๐Ÿค– AI Summary
This work addresses group fairness in policy learning, where optimal policies often induce substantial distributional shifts for sensitive subgroupsโ€”leading to implicit discrimination. To reconcile global utility maximization with subgroup distributional fairness, we propose a novel framework that jointly optimizes overall performance and subgroup-level distributional alignment. Our method explicitly regularizes subgroup outcome distributions via customizable distributional distance metrics (e.g., Wasserstein distance or KL divergence), unifying distributional objective functionals with robust constraints. It integrates empirical risk minimization and distributionally robust optimization, enabling data-driven hyperparameter selection. Theoretically, we establish regret bounds and consistency guarantees for the regularized policy class. Empirically, our approach maintains competitive overall utility while significantly reducing inter-subgroup distributional disparity and demonstrating strong generalization across diverse benchmarks.

Technology Category

Application Category

๐Ÿ“ Abstract
A decision maker typically (i) incorporates training data to learn about the relative effectiveness of the treatments, and (ii) chooses an implementation mechanism that implies an"optimal"predicted outcome distribution according to some target functional. Nevertheless, a discrimination-aware decision maker may not be satisfied achieving said optimality at the cost of heavily discriminating against subgroups of the population, in the sense that the outcome distribution in a subgroup deviates strongly from the overall optimal outcome distribution. We study a framework that allows the decision maker to penalize for such deviations, while allowing for a wide range of target functionals and discrimination measures to be employed. We establish regret and consistency guarantees for empirical success policies with data-driven tuning parameters, and provide numerical results. Furthermore, we briefly illustrate the methods in two empirical settings.
Problem

Research questions and friction points this paper is trying to address.

Balancing optimal policy learning with fairness constraints
Regularizing outcome distribution deviations across subgroups
Ensuring fairness while achieving distributional target objectives
Innovation

Methods, ideas, or system contributions that make the work stand out.

Regularizes fairness deviations in optimal policies
Employs wide range target functionals fairness measures
Establishes regret consistency guarantees empirical policies