gradient-boosted risk modeling

Designs and trains gradient-boosted tree models (commonly using XGBoost) to predict and quantify event or outcome risk, producing calibrated probabilistic risk scores. Work includes featurizing inputs (numeric, categorical, temporal, and textual attributes), encoding variables, optimizing hyperparameters with cross‑validation, and interpreting model behavior and feature importances (e.g., with SHAP) for explanation and feature selection.

gradient-boostedriskmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This paper systematically evaluates ten gradient boosting algorithms—including GBM, XGBoost, LightGBM, CatBoost, EGBM, PGBM, XGBoostLSS, and NGBoost—for insurance claim frequency and severity prediction. Using five heterogeneous public datasets, it conducts a unified benchmarking study across computational efficiency, predictive accuracy, and model interpretability. Methodologically, the work introduces: (1) the first unified evaluation framework integrating point estimation and probabilistic modeling; (2) a novel boosting mechanism embedding exposure as a bias term in frequency models; and (3) empirical validation that model fidelity and high predictive accuracy are simultaneously attainable. Results show LightGBM and XGBoostLSS achieve superior computational efficiency; CatBoost significantly enhances generalization on high-cardinality categorical features; and EGBM attains black-box-level accuracy while retaining full structural interpretability. The study thus bridges practical scalability, statistical rigor, and transparency requirements in actuarial modeling.

Compare gradient boosting algorithms for claim predictionEvaluate computational efficiency and predictive performanceHandle exposure-to-risk in frequency models with boosting

Traditional risk scores rely on linear models, which struggle to capture nonlinear effects and often produce lengthy rule sets that compromise interpretability and practical utility. This work proposes a novel approach that, for the first time, integrates gradient boosting into risk score construction, achieving both strong nonlinear modeling capacity and a human-computable, concise structure amenable to classification, regression, and time-to-event tasks. Implemented in C++ with Python and R interfaces, the method is evaluated across 12 tabular datasets. Results demonstrate that it reduces the number of rules by an average of 60% in classification tasks and 16% in time-to-event tasks, while maintaining predictive performance on par with state-of-the-art methods.

compact modelsgradient boostinginterpretability

FPBoost: Fully Parametric Gradient Boosting for Survival Analysis

Sep 20, 2024
AA
Alberto Archetti
🏛️ Politecnico di Milano | Politecnico di Torino | Chattermill

To address the overly restrictive proportional hazards assumption and the trade-off between model flexibility and interpretability in few-shot time-to-event modeling, this paper proposes PGB-Surv—the first fully parametric gradient-boosted survival model. PGB-Surv abandons the Cox assumption and instead employs decision trees to jointly optimize parameters of parametric survival distributions (e.g., Weibull, Log-Normal) for hazard function estimation, constructing flexible hazard surfaces via weighted ensemble learning. We theoretically prove that PGB-Surv is a universal approximator for any smooth hazard function. Integrated within a maximum survival likelihood estimation framework and boosted via gradient descent, the method achieves significant improvements across multiple benchmark datasets: average C-index gains of +3.2% and enhanced calibration performance. PGB-Surv thus delivers both strong generalization capability and statistically grounded interpretability—enabling reliable risk prediction and transparent parameter-level inference in low-data regimes.

Gradient BoostingInterpretable ModelsSurvival Analysis

Learning accurate and interpretable tree-based models

May 24, 2024
MB
Maria-Florina Balcan
🏛️ Carnegie Mellon University

In same-domain repeated access scenarios, decision trees and their ensembles (e.g., random forests, GBDTs) struggle to jointly optimize accuracy and interpretability. Method: This paper proposes a learnable, tunable unified tree framework. It introduces (1) a parameterized splitting criterion that continuously interpolates between entropy and Gini impurity, enabling data-adaptive optimal splits; (2) a theoretical characterization of sample complexity, providing generalization guarantees for the interpretability–accuracy trade-off; and (3) joint optimization of Bayesian decision trees, minimum-cost-complexity pruning, and ensemble hyperparameters. Results: Extensive experiments on real-world datasets demonstrate that the framework significantly improves the consistency between predictive accuracy and model interpretability. It offers both theoretical rigor—via provable generalization bounds—and practical utility—through end-to-end differentiability and seamless integration into existing tree-based pipelines. The approach bridges a critical gap between statistical performance and human-understandable structure in tree learning.

Develop data-specific tree-based learning algorithmsOptimize explainability versus accuracy trade-offTune hyperparameters in pruning and ensembles

Latest Papers

What's happening recently
View more

This work proposes a gradient boosting framework tailored for vector-valued outputs, addressing the limitations of conventional approaches that rely on dimension-wise updates or diagonal Hessian approximations and thus fail to capture interdependencies among output dimensions. By incorporating histogram-accelerated decision trees capable of supporting non-diagonal Hessian approximations and employing vector-valued leaf nodes, the proposed method enables joint modeling of the output structure. This approach relaxes the simplifying assumptions commonly imposed on vector targets in existing algorithms, substantially enhancing model expressiveness and predictive accuracy while preserving computational efficiency during training.

decision tree ensemblesgradient boostingmultinomial logistic regression

This work addresses the challenge of efficiently generating mixed-type tabular data by proposing XGenBoost, the first framework to effectively adapt XGBoost for generative modeling. The approach comprises two models: an XGBoost-driven denoising diffusion implicit model (DDIM) tailored for small datasets and a hierarchical autoregressive model designed for large-scale data. A key innovation is a Gaussian–multinomial joint diffusion mechanism that operates without one-hot encoding, combined with empirical quantile function-based dequantization and hierarchical classifiers to preserve the ordinal structure of numerical features. Evaluated across multiple benchmarks, XGenBoost consistently outperforms existing neural and tree-based generative models in generation quality while substantially reducing training costs.

generative modelingmixed-type datatabular data synthesis

This work addresses the computational inefficiency of traditional gradient boosting in multi-output prediction tasks—such as multiple quantile regression—where a separate base model must be trained for each output. The authors propose a general and efficient parallel gradient boosting algorithm that shares a unified descent direction across all outputs, requiring only a single base model per iteration and thereby substantially reducing computational overhead. Their approach overcomes existing limitations on loss functions and base learner types, supporting arbitrary combinations, and establishes the first scalable framework for multi-output conditional distribution estimation. Experiments demonstrate that the method achieves predictive accuracy comparable to XGBoost while accelerating training by several orders of magnitude, and it outperforms current nonparametric and semiparametric methods in high-dimensional settings with mixed or missing covariates.

computational efficiencyconditional distribution estimationgradient boosting

This work addresses the challenge of validating calibration in regression models—particularly critical in applications like insurance pricing where cross-subsidization across groups must be avoided—under limited and noisy data. It introduces, for the first time, gradient boosting trees into the calibration testing framework and proposes a nonparametric regression-based statistical inference method that leverages expected consistency checks to assess whether a model satisfies necessary conditions for both calibration and self-calibration. The approach provides a practical diagnostic tool for high-dimensional, nonlinear models and demonstrates high statistical power on large-scale insurance datasets, effectively identifying miscalibrated models.

auto-calibrationboosting treescalibration

Hot Scholars

QS

Qiuzhuang Sun

University of Sydney
Reliability engineeringIndustrial statisticsMaintenance
XY

Xiaohao Yang

Google
Pair Distribution FunctionX-rayDiffraction
LK

Lyndon Kennedy

Apple
machine learninginformation retrievalrecommendationsearch
MV

Maarten van Smeden

University Medical Center Utrecht
Medical statisticsMethodsPredictionData science
HH

Houman Homayoun

University of California Davis
Applied Machine LearningSystem SecurityHardware SecurityComputer Architecture