🤖 AI Summary
Traditional average treatment effects fail to capture individual heterogeneity, and under high-dimensional settings, the sublevel set structure of the conditional average treatment effect (CATE) function is complex, lacking a concise global measure of heterogeneity. This work formalizes the probability curve of CATE sublevel sets as a target parameter for the first time, revealing its non-pathwise differentiability. By integrating Grenander-type monotone estimation with debiased machine learning techniques, the authors develop a nonparametric inference framework. The proposed estimator demonstrates strong finite-sample performance and is applied empirically to randomized trial data on diabetes medication, effectively uncovering heterogeneous treatment effects across subpopulations.
📝 Abstract
The average treatment effect can obscure important heterogeneity when individuals respond differently to a treatment. While the conditional average treatment effect (CATE) function captures such heterogeneity, it is difficult to communicate when it depends on many covariates. Sublevels sets of a multivariate CATE function are equally complicated objects, but the probability of a sublevel set of a CATE function is a single number with a simple interpretation as the proportion of individuals whose expected treatment effect does not exceed a prespecified threshold. By varying the threshold, a univariate monotone curve appears which can be used to visualize the overall type and degree of heterogeneity in a population. We formalize this curve as a target parameter and show that it is not pathwise differentiable under a nonparametric model. To address this nonstandard estimation problem, we leverage recent advances in monotone function estimation and develop a Grenander-type estimator that incorporates machine learning. We also show that the best piecewise linear approximation to the curve of interest is a pathwise differentiable parameter, and we develop a debiased machine learning estimator of this approximation. We investigate our proposed estimators' finite sample performance in a sequence of numerical studies based on data synthesized from a randomized trial. The methods are illustrated in data from a randomized trial on diabetes medication.