Double Descent in Gradient Boosting Decision Trees via Split-Candidate Scaling

📅 2026-08-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the absence of a well-defined, single-axis capacity parameter for systematically studying double descent in gradient boosting decision trees (GBDT). It establishes, for the first time, the number of split candidates as a key capacity control parameter for GBDT. Through split candidate scaling, feature quantization grid analysis, boosting path dictionary modeling, and empirical tree kernel diagnostics, the study reveals that double descent arises from the interplay between the geometric structure induced by split candidates and the boosting dynamics. A characteristic double-descent pattern—where test error first increases and then decreases with the number of split candidates—is consistently observed across XGBoost, LightGBM, and CatBoost, corroborating theoretical predictions. In contrast, under the same setting, random forests exhibit only monotonic improvement without double descent.
📝 Abstract
Double descent is commonly studied by scaling an explicit capacity parameter, such as neural-network width. For gradient boosting decision trees (GBDTs), however, an analogous single-axis capacity parameter has not been established. We propose the number of split candidates as an operational capacity parameter for GBDTs. Holding other training controls fixed, increasing the split-candidate budget refines the feature-quantization grid and expands the dictionary of root-to-leaf paths from which boosting selects its updates. To analyze this expansion, we construct an empirical tree-kernel diagnostic that summarizes how candidate-induced paths group the training examples. A regime in which the empirical kernel rank grows toward the sample size and very small positive eigenvalues emerge exposes noise-sensitive directions; in this regime, test error peaks before decreasing again at larger split-candidate budgets. This perspective predicts that deeper trees should reach the regime with fewer split candidates, larger training sets should require finer grids, and label noise should make the peak more pronounced. Experiments support these predictions and show test-error peaks at intermediate split-candidate budgets across XGBoost, LightGBM, and CatBoost, whereas a random-forest control improves monotonically under the same split-candidate sweep. Taken together, our analysis and experiments support split-candidate scaling as a single-axis capacity intervention for studying GBDTs and suggest that the observed double descent arises from an interaction between candidate-induced geometry and boosting dynamics.
Problem

Research questions and friction points this paper is trying to address.

double descent
gradient boosting decision trees
capacity parameter
split candidates
feature quantization
Innovation

Methods, ideas, or system contributions that make the work stand out.

split-candidate scaling
gradient boosting decision trees
double descent
empirical tree kernel
capacity parameter
🔎 Similar Papers
No similar papers found.
R
Ryuichi Kanoh
The University of Electro-Communications; National Institute of Informatics