๐ค AI Summary
This study addresses the challenge of modeling sparse, irregularly sampled, and nonequidistant longitudinal response trajectoriesโa setting where existing methods often focus solely on scalar endpoints and neglect the full trajectory information. To overcome this limitation, we propose the Longitudinal Random Forest (LRF) framework, which integrates tree-based ensemble learning with adaptive within-node trajectory estimation to simultaneously capture individual trajectory dynamics, within-node correlations, between-node heterogeneity, and nonlinear as well as interaction effects of covariates. Key innovations include a trajectory-aware splitting criterion, two node-smoothing strategies (LRF-PACE and LRF-adaptiveLMM), mechanisms for trajectory prediction and future extrapolation, and novel metrics for variable importance and interaction frequency. Empirical results demonstrate that LRF substantially outperforms current approaches under high sparsity and effectively supports five critical clinical inference tasks.
๐ Abstract
Longitudinal studies often collect data at sparse, irregular, and unequally spaced time points. Such heterogeneity is often driven by subject-specific covariates, yet existing methods have been restricted to a scalar endpoint value, completely neglecting the underlying response trajectories. We propose a novel Longitudinal Random Forest (LRF) framework that leverages tree-based ensemble machine learning with adaptive node-wise longitudinal trajectory estimation. The LRF framework makes five methodological contributions. it captures each subject's individual response trajectory while simultaneously accommodating within-node correlation, between-node heterogeneity, and nonlinear and interactive covariate effects. It introduces a novel trajectory-based splitting criterion that maximizes trajectory separation while incorporating a size-weighted penalty; it provides two variants, Principal Analysis by Conditional Expectation (LRF-PACE) and adaptive linear mixed-effects models (LRF-adaptiveLMM), which employ nonparametric and semiparametric node-wise smoothers, respectively, while learning covariate effects in a data-driven manner. It provides a comprehensive interpretation of covariates using both the classical trajectory-based permutation variable importance measure (PVIM) and a newly proposed finite-way interaction frequency count, and it not only predicts entire trajectories for new subjects but also forecasts future trajectories for existing subjects. Extensive simulation studies demonstrate that LRF achieves superior performance over several competing methods, even under severe sparsity. The practical significance of the LRF framework lies in its ability to address five important clinical questions.