Regularisation of CART trees by summation of $p$-values

📅 2025-05-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Conventional complexity control in CART regression trees relies on stochastic cross-validation, yielding non-deterministic and irreproducible results. Method: We propose a deterministic, in-sample pruning method based on node-level statistical testing: modeling tree splitting as a multidimensional change-point detection problem and constructing a per-node p-value stopping criterion; we further derive the first theoretical upper bound on the sum of p-values to guarantee global significance control. Contribution/Results: This is the first tree-growth termination mechanism that is fully deterministic, interpretable, and applicable to arbitrary-dimensional covariates; it naturally extends to automatic early stopping in boosting. Theoretically, it achieves high detection power for non-weak signals. Empirical evaluations on synthetic and real-world datasets confirm its superior generalization performance and stability. In boosting, it enables root-node-triggered deterministic termination, substantially enhancing reproducibility and robustness.

Technology Category

Machine Learning: Ensemble MethodsReasoning under Uncertainty: Stochastic OptimizationConstraint Satisfaction and Optimization: Satisfiability Modulo Theories

Application Category

Web Mining and Content Analysis: Robustness and generalizability of Web mining methodsGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsSemantics and Knowledge: Methods, algorithms and applications for the development of semantic models, knowledge graphs and other forms of structured data models with machine-interpretable semantics
📝 Abstract
The standard procedure to decide on the complexity of a CART regression tree is to use cross-validation with the aim of obtaining a predictor that generalises well to unseen data. The randomness in the selection of folds implies that the selected CART tree is not a deterministic function of the data. We propose a deterministic in-sample method that can be used for stopping the growing of a CART tree based on node-wise statistical tests. This testing procedure is derived using a connection to change point detection, where the null hypothesis corresponds to that there is no signal. The suggested $p$-value based procedure allows us to consider covariate vectors of arbitrary dimension and allows us to bound the $p$-value of an entire tree from above. Further, we show that the test detects a not-too-weak signal with a high probability, given a not-too-small sample size. We illustrate our methodology and the asymptotic results on both simulated and real world data. Additionally, we illustrate how our $p$-value based method can be used as an automatic deterministic early stopping procedure for tree-based boosting. The boosting iterations stop when the tree to be added consists only of a root node.
Problem

Research questions and friction points this paper is trying to address.

Deterministic in-sample method for stopping CART tree growth
Node-wise statistical tests based on p-values for tree complexity
Automatic early stopping for tree-based boosting using p-values
Innovation

Methods, ideas, or system contributions that make the work stand out.

Deterministic in-sample method for CART trees
Node-wise statistical tests for stopping tree growth
P-value based early stopping for tree boosting
💼 Related Jobs
No related jobs found.
N
Nils Engler
Department of Mathematics, Stockholm University, Sweden
Mathias Lindholm
Mathias Lindholm
Senior lecturer, Dep. Math., Div. Math. Statistics, Stockholm University
Applied probabilitystatistics
F
Filip Lindskog
Department of Mathematics, Stockholm University, Sweden
T
Taariq Nazar
Department of Mathematics, Stockholm University, Sweden