TreeWalker: Partial Evaluation for Grouped Tree-Ensemble Inference

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the redundant computation caused by shared features during grouped inference in tree ensemble models by proposing an optimization framework based on partial evaluation. The method decouples feature static and dynamic components, leveraging bitmask partitioning to enable single-pass inference. Furthermore, it introduces a theory of structured work decomposition, proving that the per-row computational cost asymptotically approaches the theoretical lower bound as group size increases. The proposed system maintains full compatibility with standard LightGBM and XGBoost models. Experimental results demonstrate speedups ranging from 2.5× to 7.8× on Intel platforms, with even more pronounced gains on Arm architectures, while preserving floating-point precision comparable to native implementations.
📝 Abstract
Many inference workloads evaluate a trained tree ensemble on row groups that share feature values: discrete-time survival models expand each patient into $G$ time steps, click-through-rate models score every item in a search session, and scenario analyses vary a few inputs while holding the rest fixed. Standard inference treats each row independently and repeats the shared work $G$ times. We present TreeWalker, which applies partial evaluation to grouped inference: constant features are static, varying features dynamic. It walks each tree once per group, partitions a row bitmask at varying splits, and skips empty subtrees. Training is unchanged: TreeWalker reads standard LightGBM and XGBoost models. We prove a structural work decomposition: per-tree work splits into the constant-projected subtree size $|T_c|$, $G$ leaf writes, and a predicate-mask provisioning cost $Q$. For the trace evaluator, per-row work approaches a $(d_v+1)/(d+1)$ fraction of a row-independent walk as $G \to \infty$. On Intel, TreeWalker is 2.5-3.2$\times$ faster than a row-independent traversal at the reference configuration ($T=500$, $L=8$) and 6.8-7.8$\times$ faster at $G=128$ on the survival datasets, with larger gains on Arm. On a scenario-analysis benchmark it is faster in all 16 configurations on both architectures. For f64 models, outputs match treelite's GTIL up to summation order; for f32 models, f64 accumulation is closer to a Kahan-compensated reference than native f32 on 99.98% of rows and never farther.
Problem

Research questions and friction points this paper is trying to address.

tree-ensemble inference
grouped inference
partial evaluation
redundant computation
shared features
Innovation

Methods, ideas, or system contributions that make the work stand out.

Partial Evaluation
Grouped Inference
Tree Ensemble
Bitmask Partitioning
Work Decomposition
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Durmus Karatay
Shopify
R
Richard Newman
Shopify