Score
Designs and implements gradient-boosted decision tree ensembles whose leaf outputs are vectors, producing multiclass or multivariate predictions. This includes deriving and applying vector-form gradients and Hessians for multivariate loss functions, optimizing vector-valued objectives directly, and fitting tree ensembles that produce vector predictions.
This work proposes a gradient boosting framework tailored for vector-valued outputs, addressing the limitations of conventional approaches that rely on dimension-wise updates or diagonal Hessian approximations and thus fail to capture interdependencies among output dimensions. By incorporating histogram-accelerated decision trees capable of supporting non-diagonal Hessian approximations and employing vector-valued leaf nodes, the proposed method enables joint modeling of the output structure. This approach relaxes the simplifying assumptions commonly imposed on vector targets in existing algorithms, substantially enhancing model expressiveness and predictive accuracy while preserving computational efficiency during training.
Tree ensemble models, such as random forests and gradient-boosted trees, often suffer from limited interpretability, undermining user trust. This work proposes the first rigorous logical explanation framework tailored specifically for tree ensembles, integrating formal verification techniques with intrinsic structural properties of decision trees to construct explanations that precisely capture the model’s actual decision behavior. The approach formally guarantees correctness and faithfulness of the generated explanations, yielding verifiable and high-fidelity justifications for individual predictions. By ensuring that explanations are both logically sound and aligned with the model’s true reasoning process, the method significantly enhances the transparency and credibility of tree-based ensemble models.
In same-domain repeated access scenarios, decision trees and their ensembles (e.g., random forests, GBDTs) struggle to jointly optimize accuracy and interpretability. Method: This paper proposes a learnable, tunable unified tree framework. It introduces (1) a parameterized splitting criterion that continuously interpolates between entropy and Gini impurity, enabling data-adaptive optimal splits; (2) a theoretical characterization of sample complexity, providing generalization guarantees for the interpretability–accuracy trade-off; and (3) joint optimization of Bayesian decision trees, minimum-cost-complexity pruning, and ensemble hyperparameters. Results: Extensive experiments on real-world datasets demonstrate that the framework significantly improves the consistency between predictive accuracy and model interpretability. It offers both theoretical rigor—via provable generalization bounds—and practical utility—through end-to-end differentiability and seamless integration into existing tree-based pipelines. The approach bridges a critical gap between statistical performance and human-understandable structure in tree learning.
Embedding decision trees—including ensembles—into optimization problems suffers from low modeling accuracy and poor computational efficiency due to weak linear relaxations in existing mixed-integer programming (MIP) formulations. Method: We propose an ideal MIP formulation based on the union of projection polytopes, explicitly capturing tree logic via binary feature representations and extending to one-dimensional continuous features. Contribution/Results: We prove, for the first time under binary feature encoding, that allowing repeated splits on the same feature eliminates fractional extreme points in the linear relaxation. We further derive the ideal MIP characterization for univariate continuous features. Our formulation substantially tightens the linear relaxation and reduces the number of extreme points in the feasible region. On low-dimensional feature instances, average solution time decreases by an order of magnitude. This advancement significantly improves both the embeddability of tree models into optimization frameworks and their computational scalability.
This paper addresses scalar regression and tensor-to-tensor regression tasks with tensor-valued inputs. We propose the Tensor-Input Tree (TT) model—a novel decision tree architecture natively supporting tensor inputs. TT is the first additive tree ensemble framework extended to tensor–tensor regression and introduces efficient randomized and deterministic splitting algorithms that jointly ensure theoretical interpretability and computational efficiency. A rigorous theoretical analysis establishes a generalization error bound for TT. Empirical evaluation on both synthetic and real-world datasets demonstrates that TT achieves 2–4 orders of magnitude faster training than tensor-input Gaussian processes and other baselines, while maintaining comparable or superior prediction accuracy. The core contributions are: (i) a tensor-adapted decision tree structure; (ii) an additive ensemble framework for tensor regression; and (iii) scalable, efficient learning algorithms with provable guarantees.
Tree ensemble methods—such as random forests and gradient boosting—exhibit strong generalization yet lack a unified theoretical foundation grounded in functional analysis. Method: This work establishes the first Reproducing Kernel Hilbert Space (RKHS) framework for tree ensembles. It constructs a data-dependent random forest kernel, rigorously proving its boundedness, continuity, and universality; formulates random forests as regularized empirical risk minimization in RKHS; and models continuous-time gradient boosting as a dynamical system on a Hilbert manifold. It further introduces Geometric Variable Importance (GVI), a kernel-geometry-based feature attribution criterion, and a kernel PCA-based interpretability method. Contribution/Results: The framework provides a variational principle unifying ensemble learning, offers a geometric interpretation of tree-based prediction, and delivers novel, theoretically grounded tools for model interpretability—thereby bridging functional analysis, statistical learning theory, and practical machine learning.
To address the inherent trade-off between high estimation variance in decision trees and the lack of interpretability in random forests, this paper proposes a novel, interpretable embedding method based on leaf-node means. Specifically, it constructs a continuous embedding space using the leaf partitions of classification trees, mapping each input to the mean feature vector of its corresponding leaf region—thereby substantially reducing estimation variance. The method further integrates bootstrap aggregation with linear discriminant analysis (LDA) to jointly optimize predictive accuracy, computational efficiency, and model transparency. Theoretically grounded and empirically validated, the approach achieves classification accuracy on par with or surpassing that of random forests and shallow neural networks across multiple synthetic and real-world benchmark datasets. Moreover, it offers significantly faster training, strong generalization, low computational overhead, and explicit, semantically meaningful decision rules—effectively reconciling fidelity, efficiency, and interpretability in tree-based learning.
Tree ensemble models, such as random forests and gradient boosting machines, suffer from limited theoretical understanding and high deployment costs. This work addresses these challenges through a spectral perspective, leveraging the eigenvalue decay of kernel operators to characterize their statistical convergence rates and establishing, for the first time, minimax-optimal convergence bounds for random forest regression. Building upon dominant eigenfunctions or singular vectors, the authors propose a general spectral distillation framework that compresses models by several orders of magnitude while preserving predictive performance. Experimental results demonstrate that this approach substantially outperforms existing pruning and rule extraction techniques, offering both theoretical rigor and practical efficiency.
This paper addresses the poor interpretability of boosted tree models by proposing an error-correction-based decision tree distillation method. Unlike conventional approaches requiring access to original training data, our method constructs a learnable correction module over the residuals of the boosted tree’s predictions and jointly optimizes both the decision tree structure and leaf-node prediction values, enabling efficient knowledge transfer from the ensemble to a single interpretable tree. Compared to baseline distillation methods such as direct retraining, our approach achieves significantly higher predictive accuracy—averaging a 3.2% AUC improvement across multiple benchmark datasets—while preserving full model interpretability. Key contributions include: (1) a differentiable residual correction mechanism that retains the nonlinear modeling capacity of complex ensembles; and (2) an end-to-end framework for joint tree architecture search and parameter optimization. Experiments demonstrate superior trade-offs between accuracy and interpretability.
该研究通过将梯度提升集成模型的叶节点值视为坐标,实现了精确对比解释,并基于此方法构建了补救措施,在五个表格数据集上进行了评估。