🤖 AI Summary
Tree ensemble methods—such as random forests and gradient boosting—exhibit strong generalization yet lack a unified theoretical foundation grounded in functional analysis.
Method: This work establishes the first Reproducing Kernel Hilbert Space (RKHS) framework for tree ensembles. It constructs a data-dependent random forest kernel, rigorously proving its boundedness, continuity, and universality; formulates random forests as regularized empirical risk minimization in RKHS; and models continuous-time gradient boosting as a dynamical system on a Hilbert manifold. It further introduces Geometric Variable Importance (GVI), a kernel-geometry-based feature attribution criterion, and a kernel PCA-based interpretability method.
Contribution/Results: The framework provides a variational principle unifying ensemble learning, offers a geometric interpretation of tree-based prediction, and delivers novel, theoretically grounded tools for model interpretability—thereby bridging functional analysis, statistical learning theory, and practical machine learning.
📝 Abstract
Random Forests and Gradient Boosting are among the most effective algorithms for supervised learning on tabular data. Both belong to the class of tree-based ensemble methods, where predictions are obtained by aggregating many randomized regression trees. In this paper, we develop a theoretical framework for analyzing such methods through Reproducing Kernel Hilbert Spaces (RKHSs) constructed on tree ensembles -- more precisely, on the random partitions generated by randomized regression trees. We establish fundamental analytical properties of the resulting Random Forest kernel, including boundedness, continuity, and universality, and show that a Random Forest predictor can be characterized as the unique minimizer of a penalized empirical risk functional in this RKHS, providing a variational interpretation of ensemble learning. We further extend this perspective to the continuous-time formulation of Gradient Boosting introduced by Dombry and Duchamps, and demonstrate that it corresponds to a gradient flow on a Hilbert manifold induced by the Random Forest RKHS. A key feature of this framework is that both the kernel and the RKHS geometry are data-dependent, offering a theoretical explanation for the strong empirical performance of tree-based ensembles. Finally, we illustrate the practical potential of this approach by introducing a kernel principal component analysis built on the Random Forest kernel, which enhances the interpretability of ensemble models, as well as GVI, a new geometric variable importance criterion.