Score
Designs, builds, and evaluates predictive models composed of decision-tree learners and their ensembles—particularly gradient-boosted trees (GBDT/XGBoost)—and constructs stacked ensembles that combine multiple tree-based models to improve predictive performance. Work includes training and hyperparameter tuning of boosted trees, constructing stacking/blending pipelines with base tree models and meta-learners, performing cross-validation and model selection, and analyzing model behavior (e.g., feature importance and prediction diagnostics).
In same-domain repeated access scenarios, decision trees and their ensembles (e.g., random forests, GBDTs) struggle to jointly optimize accuracy and interpretability. Method: This paper proposes a learnable, tunable unified tree framework. It introduces (1) a parameterized splitting criterion that continuously interpolates between entropy and Gini impurity, enabling data-adaptive optimal splits; (2) a theoretical characterization of sample complexity, providing generalization guarantees for the interpretability–accuracy trade-off; and (3) joint optimization of Bayesian decision trees, minimum-cost-complexity pruning, and ensemble hyperparameters. Results: Extensive experiments on real-world datasets demonstrate that the framework significantly improves the consistency between predictive accuracy and model interpretability. It offers both theoretical rigor—via provable generalization bounds—and practical utility—through end-to-end differentiability and seamless integration into existing tree-based pipelines. The approach bridges a critical gap between statistical performance and human-understandable structure in tree learning.
Tree ensemble models, such as random forests and gradient-boosted trees, often suffer from limited interpretability, undermining user trust. This work proposes the first rigorous logical explanation framework tailored specifically for tree ensembles, integrating formal verification techniques with intrinsic structural properties of decision trees to construct explanations that precisely capture the model’s actual decision behavior. The approach formally guarantees correctness and faithfulness of the generated explanations, yielding verifiable and high-fidelity justifications for individual predictions. By ensuring that explanations are both logically sound and aligned with the model’s true reasoning process, the method significantly enhances the transparency and credibility of tree-based ensemble models.
This work addresses the rigidity in model selection and poor interpretability inherent in conventional XGBoost–neural network ensembles. We propose an adaptive fusion framework grounded in dynamic meta-learning, which jointly leverages uncertainty quantification and feature importance as dual control signals to guide fine-grained scheduling and weighted integration of the two base models at inference time. Our key innovation lies in co-modeling uncertainty estimates and interpretability-aware metrics—specifically, feature importance—within the meta-learner’s decision process, thereby simultaneously enhancing predictive performance and decision transparency. Extensive experiments across multiple benchmark datasets demonstrate that our method consistently outperforms static ensembles and individual baselines, achieving average accuracy gains of 2.1–4.7 percentage points. Moreover, it provides auditable, instance-level rationale for model selection and feature-level attribution, supporting both reliability assessment and human-understandable explanations.
This study addresses diagnostic tasks on medical tabular data by systematically evaluating gradient-boosted decision tree (GBDT) models—including XGBoost, CatBoost, and LightGBM—in terms of predictive performance and practical deployability. We conduct the first comprehensive benchmark across multiple public medical datasets, comparing GBDTs against classical machine learning methods (SVM, logistic regression) and state-of-the-art deep tabular models (TabNet, TabTransformer). Results demonstrate that GBDT models achieve the highest average ranking, significantly outperforming all baselines in diagnostic accuracy while reducing training time by 40–75% and memory consumption by approximately 60%. These findings establish GBDTs as the method of choice for high-accuracy, low-compute medical diagnosis—offering a computationally efficient, robust, and clinically viable modeling paradigm for decision support systems.
To address the insufficient interpretability, fairness, and robustness of traditional decision trees in high-stakes prediction and prescriptive decision-making, this paper proposes a unified modeling framework based on mixed-integer optimization (MIO) and releases ODTlearn, an open-source Python package. The framework systematically integrates four classes of optimal decision trees—classification, fair classification, distributionally robust classification, and observational-data-driven prescriptive trees—enabling multi-objective trade-offs and constraint-based modeling. Designed with object-oriented principles, it supports commercial (e.g., Gurobi) and open-source (e.g., COIN-OR CBC) solvers, balancing computational efficiency and scalability. Comprehensive documentation, tutorials, and fully reproducible code are provided. Empirical evaluations demonstrate that the approach preserves strong interpretability while significantly improving fairness, out-of-distribution generalization, and individualized prescription quality in high-risk settings.
This work proposes a sequential testing–based early-stopping strategy for binary ensemble classifiers to reduce inference overhead while strictly bounding the divergence rate from predictions of the full ensemble. The approach terminates evaluation as soon as a decisive majority emerges during the sequential assessment of base models. Under three optimality criteria, the strategy can be formulated as a linear programming problem, enabling efficient computation of the optimal stopping rule. Experimental results on UCI and Grinsztajn benchmark datasets demonstrate that the method achieves an average speedup exceeding 4× while consistently maintaining prediction divergence below 0.1%.
本文提出qshap,一种快速分解梯度提升树模型中R²值的方法,以量化特征对模型性能的具体贡献。
This work addresses the computational inefficiency of traditional gradient boosting in multi-output prediction tasks—such as multiple quantile regression—where a separate base model must be trained for each output. The authors propose a general and efficient parallel gradient boosting algorithm that shares a unified descent direction across all outputs, requiring only a single base model per iteration and thereby substantially reducing computational overhead. Their approach overcomes existing limitations on loss functions and base learner types, supporting arbitrary combinations, and establishes the first scalable framework for multi-output conditional distribution estimation. Experiments demonstrate that the method achieves predictive accuracy comparable to XGBoost while accelerating training by several orders of magnitude, and it outperforms current nonparametric and semiparametric methods in high-dimensional settings with mixed or missing covariates.
This work addresses the absence of a well-defined, single-axis capacity parameter for systematically studying double descent in gradient boosting decision trees (GBDT). It establishes, for the first time, the number of split candidates as a key capacity control parameter for GBDT. Through split candidate scaling, feature quantization grid analysis, boosting path dictionary modeling, and empirical tree kernel diagnostics, the study reveals that double descent arises from the interplay between the geometric structure induced by split candidates and the boosting dynamics. A characteristic double-descent pattern—where test error first increases and then decreases with the number of split candidates—is consistently observed across XGBoost, LightGBM, and CatBoost, corroborating theoretical predictions. In contrast, under the same setting, random forests exhibit only monotonic improvement without double descent.
This work proposes a gradient boosting framework tailored for vector-valued outputs, addressing the limitations of conventional approaches that rely on dimension-wise updates or diagonal Hessian approximations and thus fail to capture interdependencies among output dimensions. By incorporating histogram-accelerated decision trees capable of supporting non-diagonal Hessian approximations and employing vector-valued leaf nodes, the proposed method enables joint modeling of the output structure. This approach relaxes the simplifying assumptions commonly imposed on vector targets in existing algorithms, substantially enhancing model expressiveness and predictive accuracy while preserving computational efficiency during training.