Score
Designs, trains, and tunes gradient-boosted decision-tree ensembles (e.g., XGBoost) by specifying boosting algorithms, stagewise loss minimization, tree structures, and hyperparameters such as learning rate, tree depth, and regularization. Builds evaluation and validation workflows to measure predictive performance, assess model complexity and overfitting, and produce deployable gradient-boosted tree models.
In same-domain repeated access scenarios, decision trees and their ensembles (e.g., random forests, GBDTs) struggle to jointly optimize accuracy and interpretability. Method: This paper proposes a learnable, tunable unified tree framework. It introduces (1) a parameterized splitting criterion that continuously interpolates between entropy and Gini impurity, enabling data-adaptive optimal splits; (2) a theoretical characterization of sample complexity, providing generalization guarantees for the interpretability–accuracy trade-off; and (3) joint optimization of Bayesian decision trees, minimum-cost-complexity pruning, and ensemble hyperparameters. Results: Extensive experiments on real-world datasets demonstrate that the framework significantly improves the consistency between predictive accuracy and model interpretability. It offers both theoretical rigor—via provable generalization bounds—and practical utility—through end-to-end differentiability and seamless integration into existing tree-based pipelines. The approach bridges a critical gap between statistical performance and human-understandable structure in tree learning.
This paper systematically evaluates six linear programming (LP)-based fully-corrective boosting methods—including two newly proposed approaches, NM-Boost and QRLP-Boost—against mainstream heuristic gradient boosting frameworks (e.g., XGBoost, LightGBM) across 20 benchmark datasets. Addressing limitations in accuracy, ensemble sparsity, margin distribution, real-time inference efficiency, and hyperparameter robustness, the study provides the first large-scale empirical evidence that LP-based boosters achieve accuracy competitive with state-of-the-art gradient boosting when using shallow decision trees as base learners, while drastically reducing ensemble size (sparsity improved by multiple-fold). Crucially, these methods enable lossless sparsification of pre-trained ensembles. The results demonstrate that global optimization—central to LP-based boosting—enhances both robustness and interpretability, revealing a promising new paradigm for trustworthy ensemble learning.
This study addresses the critical dependence of tree boosting models’ generalization performance on hyperparameter configuration, a domain lacking systematic comparison and guiding principles among existing tuning methods. Conducting a large-scale empirical evaluation across 59 regression and classification datasets under a unified experimental framework, the authors compare prominent hyperparameter optimization approaches—including random search, grid search, Tree-structured Parzen Estimator (TPE), Gaussian process Bayesian optimization, Hyperband, and SMAC. The findings reveal that SMAC consistently outperforms other methods across most tasks, while default hyperparameters yield suboptimal results. Notably, all hyperparameters significantly influence performance, contradicting the common assumption that only a few are critical. Effective tuning requires extensive sampling (>100 trials), and for regression tasks, employing early stopping proves superior to including the number of iterations within the search space.
Existing tree-based models for tabular data generation lack explicit density modeling capabilities. Method: This paper proposes NRGBoost—the first energy-based generative gradient-boosting tree framework. It extends discriminative GBDT to a generative model capable of learning unnormalized densities (up to a constant) by optimizing an energy potential function via second-order Taylor expansion, where tree structure updates are driven jointly by gradients and Hessian matrices in a generative boosting scheme. Results: On multiple real-world tabular datasets, NRGBoost matches XGBoost’s discriminative performance while significantly outperforming other generative tree methods; its sampling quality rivals that of neural generative models, and it supports arbitrary conditional inference and joint sampling. The core contribution is the first realization of second-order optimization–driven generative tree boosting, unifying high discriminative accuracy with principled generative capability.
Traditional gradient boosting trees (GBTs) employ static tree structures and fixed splitting criteria, limiting their adaptability to evolving gradient distributions and task-specific阶段性 characteristics during training. To address this, we propose MorphBoost—a novel GBT framework featuring self-organizing, dynamic tree structures. Its core contributions are: (1) a gradient-statistics-driven adaptive splitting function that adjusts split decisions based on evolving gradient moments; (2) problem-fingerprint identification guided by training progress, enabling dynamic regulation of learning pressure across stages; and (3) vectorized tree prediction coupled with interaction-aware feature importance modeling. Evaluated on 10 standard benchmark datasets, MorphBoost achieves an average accuracy gain of 0.84% over XGBoost and attains state-of-the-art performance on 40% of the datasets. Moreover, it significantly reduces prediction variance and improves minimum accuracy, demonstrating superior stability and generalization capability.
This work proposes a gradient boosting framework tailored for vector-valued outputs, addressing the limitations of conventional approaches that rely on dimension-wise updates or diagonal Hessian approximations and thus fail to capture interdependencies among output dimensions. By incorporating histogram-accelerated decision trees capable of supporting non-diagonal Hessian approximations and employing vector-valued leaf nodes, the proposed method enables joint modeling of the output structure. This approach relaxes the simplifying assumptions commonly imposed on vector targets in existing algorithms, substantially enhancing model expressiveness and predictive accuracy while preserving computational efficiency during training.
This work addresses the absence of a well-defined, single-axis capacity parameter for systematically studying double descent in gradient boosting decision trees (GBDT). It establishes, for the first time, the number of split candidates as a key capacity control parameter for GBDT. Through split candidate scaling, feature quantization grid analysis, boosting path dictionary modeling, and empirical tree kernel diagnostics, the study reveals that double descent arises from the interplay between the geometric structure induced by split candidates and the boosting dynamics. A characteristic double-descent pattern—where test error first increases and then decreases with the number of split candidates—is consistently observed across XGBoost, LightGBM, and CatBoost, corroborating theoretical predictions. In contrast, under the same setting, random forests exhibit only monotonic improvement without double descent.
本文提出qshap,一种快速分解梯度提升树模型中R²值的方法,以量化特征对模型性能的具体贡献。
This work addresses the computational inefficiency of traditional gradient boosting in multi-output prediction tasks—such as multiple quantile regression—where a separate base model must be trained for each output. The authors propose a general and efficient parallel gradient boosting algorithm that shares a unified descent direction across all outputs, requiring only a single base model per iteration and thereby substantially reducing computational overhead. Their approach overcomes existing limitations on loss functions and base learner types, supporting arbitrary combinations, and establishes the first scalable framework for multi-output conditional distribution estimation. Experiments demonstrate that the method achieves predictive accuracy comparable to XGBoost while accelerating training by several orders of magnitude, and it outperforms current nonparametric and semiparametric methods in high-dimensional settings with mixed or missing covariates.
This work addresses the lack of global convergence in Newton boosting trees under general convex losses, which can lead to divergence. The authors propose a gradient-regularized Newton boosting framework that introduces an adaptive ℓ² regularization term at each iteration, proportional to the square root of the gradient norm. By integrating constrained Newton descent with Hessian-Lipschitz analysis and leveraging standard weak learner assumptions, they establish the first global convergence guarantee for Newton boosting trees. The resulting algorithm achieves a convergence rate of 𝒪(1/k²) under general convex losses—matching the rate of first-order methods equipped with Nesterov momentum. Empirical results confirm the stable convergence of the proposed method, whereas classical Newton boosting may diverge.