gradient-boosted tree model development

Designs, trains, and tunes gradient-boosted decision-tree ensembles (e.g., XGBoost) by specifying boosting algorithms, stagewise loss minimization, tree structures, and hyperparameters such as learning rate, tree depth, and regularization. Builds evaluation and validation workflows to measure predictive performance, assess model complexity and overfitting, and produce deployable gradient-boosted tree models.

gradient-boostedtreemodeldevelopment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$185K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Learning accurate and interpretable tree-based models

May 24, 2024
MB
Maria-Florina Balcan
🏛️ Carnegie Mellon University

In same-domain repeated access scenarios, decision trees and their ensembles (e.g., random forests, GBDTs) struggle to jointly optimize accuracy and interpretability. Method: This paper proposes a learnable, tunable unified tree framework. It introduces (1) a parameterized splitting criterion that continuously interpolates between entropy and Gini impurity, enabling data-adaptive optimal splits; (2) a theoretical characterization of sample complexity, providing generalization guarantees for the interpretability–accuracy trade-off; and (3) joint optimization of Bayesian decision trees, minimum-cost-complexity pruning, and ensemble hyperparameters. Results: Extensive experiments on real-world datasets demonstrate that the framework significantly improves the consistency between predictive accuracy and model interpretability. It offers both theoretical rigor—via provable generalization bounds—and practical utility—through end-to-end differentiability and seamless integration into existing tree-based pipelines. The approach bridges a critical gap between statistical performance and human-understandable structure in tree learning.

Develop data-specific tree-based learning algorithmsOptimize explainability versus accuracy trade-offTune hyperparameters in pruning and ensembles

Boosting Revisited: Benchmarking and Advancing LP-Based Ensemble Methods

Jul 24, 2025
FA
Fabian Akkerman
🏛️ University of Twente | Polytechnique Montréal | LAAS-CNRS | Université de Toulouse

This paper systematically evaluates six linear programming (LP)-based fully-corrective boosting methods—including two newly proposed approaches, NM-Boost and QRLP-Boost—against mainstream heuristic gradient boosting frameworks (e.g., XGBoost, LightGBM) across 20 benchmark datasets. Addressing limitations in accuracy, ensemble sparsity, margin distribution, real-time inference efficiency, and hyperparameter robustness, the study provides the first large-scale empirical evidence that LP-based boosters achieve accuracy competitive with state-of-the-art gradient boosting when using shallow decision trees as base learners, while drastically reducing ensemble size (sparsity improved by multiple-fold). Crucially, these methods enable lossless sparsification of pre-trained ensembles. The results demonstrate that global optimization—central to LP-based boosting—enhances both robustness and interpretability, revealing a promising new paradigm for trustworthy ensemble learning.

Analyzing ensemble sparsity and performance trade-offsComparing novel methods with state-of-the-art heuristicsEvaluating LP-based boosting methods empirically

This study addresses the critical dependence of tree boosting models’ generalization performance on hyperparameter configuration, a domain lacking systematic comparison and guiding principles among existing tuning methods. Conducting a large-scale empirical evaluation across 59 regression and classification datasets under a unified experimental framework, the authors compare prominent hyperparameter optimization approaches—including random search, grid search, Tree-structured Parzen Estimator (TPE), Gaussian process Bayesian optimization, Hyperband, and SMAC. The findings reveal that SMAC consistently outperforms other methods across most tasks, while default hyperparameters yield suboptimal results. Notably, all hyperparameters significantly influence performance, contradicting the common assumption that only a few are critical. Effective tuning requires extensive sampling (>100 trials), and for regression tasks, employing early stopping proves superior to including the number of iterations within the search space.

boosting iterationshyperparameter optimizationmodel accuracy

NRGBoost: Energy-Based Generative Boosted Trees

Oct 04, 2024
JB
Joao Bravo
🏛️ Feedzai

Existing tree-based models for tabular data generation lack explicit density modeling capabilities. Method: This paper proposes NRGBoost—the first energy-based generative gradient-boosting tree framework. It extends discriminative GBDT to a generative model capable of learning unnormalized densities (up to a constant) by optimizing an energy potential function via second-order Taylor expansion, where tree structure updates are driven jointly by gradients and Hessian matrices in a generative boosting scheme. Results: On multiple real-world tabular datasets, NRGBoost matches XGBoost’s discriminative performance while significantly outperforming other generative tree methods; its sampling quality rivals that of neural generative models, and it supports arbitrary conditional inference and joint sampling. The core contribution is the first realization of second-order optimization–driven generative tree boosting, unifying high discriminative accuracy with principled generative capability.

Competes with neural networks in sampling performanceExtends tree-based methods for generative tasksModels data density explicitly for diverse applications

MorphBoost: Self-Organizing Universal Gradient Boosting with Adaptive Tree Morphing

Nov 17, 2025
BK
Boris Kriuk
🏛️ Hong Kong University of Science and Technology

Traditional gradient boosting trees (GBTs) employ static tree structures and fixed splitting criteria, limiting their adaptability to evolving gradient distributions and task-specific阶段性 characteristics during training. To address this, we propose MorphBoost—a novel GBT framework featuring self-organizing, dynamic tree structures. Its core contributions are: (1) a gradient-statistics-driven adaptive splitting function that adjusts split decisions based on evolving gradient moments; (2) problem-fingerprint identification guided by training progress, enabling dynamic regulation of learning pressure across stages; and (3) vectorized tree prediction coupled with interaction-aware feature importance modeling. Evaluated on 10 standard benchmark datasets, MorphBoost achieves an average accuracy gain of 0.84% over XGBoost and attains state-of-the-art performance on 40% of the datasets. Moreover, it significantly reduces prediction variance and improves minimum accuracy, demonstrating superior stability and generalization capability.

Adaptive split functions adjust to problem complexity using gradient statisticsAutomatic problem fingerprinting configures parameters across diverse machine learning tasksDynamic tree morphing adapts to evolving gradient distributions during training

Latest Papers

What's happening recently
View more

This work proposes a gradient boosting framework tailored for vector-valued outputs, addressing the limitations of conventional approaches that rely on dimension-wise updates or diagonal Hessian approximations and thus fail to capture interdependencies among output dimensions. By incorporating histogram-accelerated decision trees capable of supporting non-diagonal Hessian approximations and employing vector-valued leaf nodes, the proposed method enables joint modeling of the output structure. This approach relaxes the simplifying assumptions commonly imposed on vector targets in existing algorithms, substantially enhancing model expressiveness and predictive accuracy while preserving computational efficiency during training.

decision tree ensemblesgradient boostingmultinomial logistic regression

This work addresses the absence of a well-defined, single-axis capacity parameter for systematically studying double descent in gradient boosting decision trees (GBDT). It establishes, for the first time, the number of split candidates as a key capacity control parameter for GBDT. Through split candidate scaling, feature quantization grid analysis, boosting path dictionary modeling, and empirical tree kernel diagnostics, the study reveals that double descent arises from the interplay between the geometric structure induced by split candidates and the boosting dynamics. A characteristic double-descent pattern—where test error first increases and then decreases with the number of split candidates—is consistently observed across XGBoost, LightGBM, and CatBoost, corroborating theoretical predictions. In contrast, under the same setting, random forests exhibit only monotonic improvement without double descent.

capacity parameterdouble descentfeature quantization

This work addresses the computational inefficiency of traditional gradient boosting in multi-output prediction tasks—such as multiple quantile regression—where a separate base model must be trained for each output. The authors propose a general and efficient parallel gradient boosting algorithm that shares a unified descent direction across all outputs, requiring only a single base model per iteration and thereby substantially reducing computational overhead. Their approach overcomes existing limitations on loss functions and base learner types, supporting arbitrary combinations, and establishes the first scalable framework for multi-output conditional distribution estimation. Experiments demonstrate that the method achieves predictive accuracy comparable to XGBoost while accelerating training by several orders of magnitude, and it outperforms current nonparametric and semiparametric methods in high-dimensional settings with mixed or missing covariates.

computational efficiencyconditional distribution estimationgradient boosting

This work addresses the lack of global convergence in Newton boosting trees under general convex losses, which can lead to divergence. The authors propose a gradient-regularized Newton boosting framework that introduces an adaptive ℓ² regularization term at each iteration, proportional to the square root of the gradient norm. By integrating constrained Newton descent with Hessian-Lipschitz analysis and leveraging standard weak learner assumptions, they establish the first global convergence guarantee for Newton boosting trees. The resulting algorithm achieves a convergence rate of 𝒪(1/k²) under general convex losses—matching the rate of first-order methods equipped with Nesterov momentum. Empirical results confirm the stable convergence of the proposed method, whereas classical Newton boosting may diverge.

convex optimizationGBDTglobal convergence

Hot Scholars

JF

Junyi Fan

University of Southern California
machine learning
SC

Shuheng Chen

University of Southern California
Machine LearningData SciencePredictive AnalyticsClinical Prediction
KG

Kasper Green Larsen

Professor, Computer Science, Aarhus University
Data StructuresComputational GeometryTheoretical Computer ScienceMachine Learning
EP

Elham Pishgar

Assistant professor of gasteroenterology, Iran University of Medical Science
IBD EUS