🤖 AI Summary
This study addresses the lack of robustness in standard random forests when applied to biomedical data characterized by outliers, skewness, or mixed outcomes. To this end, we propose the Wilcoxon Random Forest, which employs the Wilcoxon rank-sum statistic as its splitting criterion and optimizes tree structure by maximizing rank-based impurity reduction. This approach is invariant to monotonic transformations and naturally accommodates continuous, ordinal, and mixed outcome variables. Furthermore, conditional distribution estimation is achieved by integrating subsample aggregation with forest-weighted empirical cumulative distribution functions. Experimental results demonstrate that the proposed model substantially improves quantile estimation accuracy and calibration under skewed and heteroscedastic conditions, effectively resolving practical challenges such as the limit of detection in HIV viral load measurements.
📝 Abstract
Random forests (RF) are tree-based models that capture complex, high-order predictor interactions. In biomedical studies, ordered outcomes often have outliers, skewness, or mixtures of continuous and ordinal values (e.g., due to detection limits). Standard RF splitting targets conditional means via squared errors and is sensitive to such irregularities. We propose a Wilcoxon regression tree that selects splits by maximizing a rank-based impurity reduction, equivalent to maximizing the squared Wilcoxon rank-sum statistic. Because the criterion depends only on outcome ranks, it is invariant to monotone transformations and naturally accommodates continuous, ordinal, or mixed outcomes. We develop the Wilcoxon random forest (WRF), which aggregates Wilcoxon regression trees via subsampling and estimates conditional distributions using forest-weighted empirical CDFs. We establish consistency of the WRF distribution estimator under regularity conditions. We evaluate distributional prediction using calibration diagnostics and continuous ranked probability scores, and define an out-of-bag permutation variable-importance measure. Simulations show that the WRF performs comparably to the standard quantile regression forest when errors are symmetric and homoscedastic, and yields improved quantile estimation and calibration when outcomes are skewed or heteroscedastic. We apply the WRF to predict CD4 cell count and HIV viral load six months after antiretroviral therapy initiation in a multicenter Latin American cohort, and the WRF improves threshold probability estimation and accommodates the large mass at the viral load detection limit.