🤖 AI Summary
This study addresses the vanishing gradient problem caused by symmetry in constructive classifiers, where deepening soft decision trees leads to inherited parent-node distributions and a consequent loss of learning capacity. We systematically diagnose four structural growth strategies for soft decision trees, employing theoretical derivations and multi-dataset benchmarks to reveal the underlying zero-gradient defect mechanism. To overcome this limitation, we propose a perturbation-based remedy utilizing slight asymmetric initialization. Theoretically, we rigorously prove both the root cause of this defect and the efficacy of our proposed solution. Empirically, the remedied model achieves a 0.2% accuracy improvement, while sparse growth strategies approximate full-tree performance using only 23% of the splits. Furthermore, this work delineates the applicability boundaries of each growth strategy, offering practical guidance for constructing deeper and more efficient soft decision trees.
📝 Abstract
Constructive classifiers add structure while they train: a level to a tree, a unit to a hidden layer, a split at a leaf. This paper asks what each of four such growth decisions buys, measured under one protocol on 24 datasets, and gives an exact diagnosis and fix for the one that buys nothing. The diagnosis concerns the natural way to deepen a soft decision tree: turn every leaf into a gate whose two children inherit the parent's class distribution, so the function is unchanged. I prove that this leaves the gradient of every new gate identically zero and with the gate at 1/2, gives the two children identical gradients, so under this construction the added level can never learn. That predicts where it costs: nothing on two-class problems, where two leaves already suffice and a great deal where more classes need more leaves. Measured, the cost is -0.2 points over 13 binary datasets and 40.2 points over 7 multi-class ones, and it tracks the number of classes (Spearman 0.64), not the number of features (0.03). The fix is any perturbation of the children. Neither its size nor its direction matters: a residual-directed initialisation changes accuracy by +0.20 points against noise. The other three decisions each buy one thing. Fitting a new hidden unit to the residual before installing it buys a smaller network but not a more accurate one. Splitting one leaf at a time buys sparsity; it lost accuracy until I found that the split started its two children identical and untrained, the same defect in another place; with the children inheriting the parent and the symmetry broken, per-leaf growth comes within 2.3 points of a complete tree using 23% of its splits. Requiring statistical significance before a node gets a more expressive split buys nothing. Every number comes from the measurement scripts.