🤖 AI Summary
This study addresses the limited understanding of how well conditional average treatment effect (CATE) models can discriminate heterogeneous treatment effects across subpopulations. The authors construct 20 theoretical outcome distributions varying in both average treatment effect and heterogeneity levels, and systematically evaluate the performance of three discrimination metrics—benefit c-statistic, benefit concentration, and PAPE—under ideal CATE assumptions. Their analysis reveals marked differences in the metrics’ dependence on treatment effect heterogeneity: benefit concentration achieves perfect discrimination even under minimal heterogeneity, whereas high values of the c-statistic and PAPE require substantially stronger heterogeneity. These findings challenge the common practice of relying on a single metric to assess CATE model performance and underscore the sensitivity of metric selection to the underlying degree of treatment effect heterogeneity.
📝 Abstract
Analyzing the heterogeneity of treatment effects is crucial in personalized medicine to identify which patients will benefit from specific treatments. The performance of a conditional average treatment effects model to guide treatment decisions can be assessed in different ways, with an important one being the model's ability to effectively discriminate between individuals who benefit from the treatment and those who do not. While many methods and algorithms have been proposed to develop conditional average treatment effects models and individualized treatment rules, little is known about the discriminative ability that can be achieved according to the population's underlying distribution of treatment effects. In this work, we computed the discrimination that can be achieved under oracle CATE for a panel of 20 distributions with varying average treatment effects and levels of heterogeneity. The assessment included the following discrimination metrics: the c-statistic for benefit, the concentration of benefit, and the population average prescriptive effect (PAPE). Results showed that the three metrics employed in this study did not require the same levels of treatment effect heterogeneity to lead to high discrimination results. Notably, achieving high c-statistic for benefit and PAPE values required greater heterogeneity than obtaining high concentration of benefit values. The three metrics considered behave very differently across the distributions. For instance, the concentration of benefit can indicate perfect discrimination in settings with negligible treatment-effects heterogeneity.