Representation Multiplicity in Causal Forests

๐Ÿ“… 2026-09-10
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
่ฏฅ็ ”็ฉถ่งฃๅ†ณไบ†ๅ› ๆžœๆฃฎๆž—ไธญๅ› ๅ˜้‡็ผ–็ ๅ†—ไฝ™ๅฏผ่‡ด็š„้ข„ๆต‹ไธไธ€่‡ด้—ฎ้ข˜๏ผŒ้€š่ฟ‡ๅˆ†็ป„็”Ÿๆˆ็›ธๅŒๅˆ†ๅ‰ฒ็š„ๅ˜้‡ๆฅๆขๅค้ข„ๆต‹ไธๅ˜ๆ€งใ€‚
๐Ÿ“ Abstract
Including covariates alongside strictly monotone encodings can change a causal forest's treatment decisions without adding information. Random feature selection favors covariates represented by multiple columns. Under stated conditions, I show that this imbalance can persist as samples grow. Treatment effect components associated with other covariates are omitted, attenuated, or recovered depending on their inclusion probabilities and tree depth. Simulations and a job-training replication illustrate sensitivity to redundant encodings. Sampling groups of variables that generate identical splits restores prediction invariance on the grouping data when fitting and randomization are held fixed.
Problem

Research questions and friction points this paper is trying to address.

causal forests
covariates
monotone encodings
random feature selection
treatment decisions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Forests
Representation Multiplicity
Monotone Encodings
Feature Selection
Prediction Invariance
Y
Yi Niu
Department of Economics, University of Pennsylvania