Aggregation Trees

📅 2024-10-15
🏛️ Social Science Research Network
📈 Citations: 1
Influential: 0
📄 PDF

career value

220K/year
🤖 AI Summary
This paper addresses the challenge of subgroup identification for heterogeneous treatment effect (HTE) estimation in causal inference. We propose an interpretable, nested nonparametric subgroup partitioning tree framework that integrates honest splitting, debiased machine learning (DML), and an aggregation tree structure. To our knowledge, this is the first method ensuring subgroup nesting while enabling unbiased, asymptotically normal statistical inference for average treatment effects (ATEs) within each subgroup. Unlike existing approaches, it mitigates p-hacking risks and jointly optimizes subgroup granularity, interpretability, and statistical validity. Simulation studies demonstrate substantially improved power for detecting heterogeneity. Empirical analysis of maternal smoking’s impact on newborn birth weight reveals systematic HTE patterns driven by parental characteristics and delivery-related factors.

Technology Category

Application Category

📝 Abstract
Uncovering the heterogeneous effects of particular policies or"treatments"is a key concern for researchers and policymakers. A common approach is to report average treatment effects across subgroups based on observable covariates. However, the choice of subgroups is crucial as it poses the risk of $p$-hacking and requires balancing interpretability with granularity. This paper proposes a nonparametric approach to construct heterogeneous subgroups. The approach enables a flexible exploration of the trade-off between interpretability and the discovery of more granular heterogeneity by constructing a sequence of nested groupings, each with an optimality property. By integrating our approach with"honesty"and debiased machine learning, we provide valid inference about the average treatment effect of each group. We validate the proposed methodology through an empirical Monte-Carlo study and apply it to revisit the impact of maternal smoking on birth weight, revealing systematic heterogeneity driven by parental and birth-related characteristics.
Problem

Research questions and friction points this paper is trying to address.

Uncovering heterogeneous treatment effects across diverse population subgroups
Balancing interpretability with granularity in subgroup selection methodology
Providing valid inference for treatment effects while preventing p-hacking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Nonparametric construction of heterogeneous subgroups
Nested groupings balancing interpretability and granularity
Integration with honesty and debiased machine learning
🔎 Similar Papers