probabilistic decision tree inference

Designs, builds, or analyzes decision-tree models and inference procedures that represent and propagate probability distributions through nodes and soft branchings to produce calibrated probabilistic predictions and uncertainty estimates while preserving interpretable decision paths. Work includes specifying probabilistic routing or soft-split rules, implementing algorithms to compute posterior or predictive distributions through the tree, and evaluating calibration and interpretability of the resulting probabilistic tree models.

probabilisticdecisiontreeinference

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Efficient distributional regression trees learning algorithms for calibrated non-parametric probabilistic forecasts

Feb 07, 2025
DQ
Duchemin Quentin
🏛️ Swiss Data Science Center | EPFL | ETH Zürich

To address the growing demand for trustworthy AI, this work proposes a calibratable nonparametric probabilistic regression tree that directly models the conditional cumulative distribution function (cCDF) to produce prediction intervals with high accuracy, reliable coverage, and statistical calibration. Methodologically, it introduces efficient tree-splitting algorithms tailored to the weighted interval score (WIS) and continuous ranked probability score (CRPS)—the first such formulations for these proper scoring rules. To accelerate split search and gradient updates, it synergistically integrates min-max heaps, weight-balanced binary trees, and Fenwick trees. Additionally, the framework supports group-conditional conformal prediction, ensuring statistical validity while preserving model interpretability. Empirical evaluation demonstrates superior calibration, coverage, and computational efficiency compared to state-of-the-art baselines across diverse benchmarks. The method achieves a balanced advantage in critical scientific and engineering applications where both uncertainty quantification rigor and practical deployability are essential.

Develop efficient algorithms for probabilistic regression treesEnhance calibration and coverage in non-parametric forecastsEnsure interpretability and group-conditional coverage in predictions

Learning accurate and interpretable tree-based models

May 24, 2024
MB
Maria-Florina Balcan
🏛️ Carnegie Mellon University

In same-domain repeated access scenarios, decision trees and their ensembles (e.g., random forests, GBDTs) struggle to jointly optimize accuracy and interpretability. Method: This paper proposes a learnable, tunable unified tree framework. It introduces (1) a parameterized splitting criterion that continuously interpolates between entropy and Gini impurity, enabling data-adaptive optimal splits; (2) a theoretical characterization of sample complexity, providing generalization guarantees for the interpretability–accuracy trade-off; and (3) joint optimization of Bayesian decision trees, minimum-cost-complexity pruning, and ensemble hyperparameters. Results: Extensive experiments on real-world datasets demonstrate that the framework significantly improves the consistency between predictive accuracy and model interpretability. It offers both theoretical rigor—via provable generalization bounds—and practical utility—through end-to-end differentiability and seamless integration into existing tree-based pipelines. The approach bridges a critical gap between statistical performance and human-understandable structure in tree learning.

Develop data-specific tree-based learning algorithmsOptimize explainability versus accuracy trade-offTune hyperparameters in pruning and ensembles

What Are the Odds? Improving the foundations of Statistical Model Checking

Apr 08, 2024
TM
Tobias Meggendorfer
🏛️ Lancaster University Leipzig | Institute of Science and Technology Austria | Dresden University of Technology

For Markov decision processes (MDPs) with unknown transition probabilities, existing statistical model checking (SMC) algorithms suffer from high sample complexity and weak theoretical guarantees. Method: We introduce tight concentration inequalities—specifically, the Bretagnolle–Huber and Empirical Bernstein bounds—into the SMC framework for the first time, and design adaptive, structure-aware statistical estimators that exploit MDP topology. Contribution/Results: Theoretically, our approach yields significantly tighter and more general probably approximately correct (PAC) guarantees. Empirically, it reduces required sample sizes by up to two orders of magnitude on standard verification benchmarks. This work establishes a new paradigm for efficient and reliable formal verification of uncertain systems.

Enhancing accuracy of transition probability estimationImproving statistical methods for model checking MDPsReducing sample size requirements for SMC algorithms

Approximate Bayesian Computation sequential Monte Carlo via random forests

Jun 22, 2024
KN
Khanh N. Dinh
🏛️ Columbia University | Ecole Nationale des Ponts et Chaussées | University of Southern California | Stony Brook University

In likelihood-intractable scenarios, Approximate Bayesian Computation (ABC) faces key bottlenecks: difficulty in selecting informative summary statistics, complex tuning of distance metrics and tolerance thresholds, high computational cost, and substantial posterior uncertainty due to weak prior information. To address these challenges, this paper proposes ABC-SMC-(D)RF—a novel method integrating Distributed Random Forests (DRF) into the Sequential Monte Carlo (SMC) framework for the first time. It enables end-to-end learning of discriminative features, eliminating manual design of summary statistics and distance functions. Coupled with adaptive tolerance scheduling and a pre-acceptance shrinkage mechanism, it iteratively concentrates sampling on high-probability parameter regions. Experiments across diverse deterministic and stochastic models demonstrate significant improvements in posterior estimation accuracy and robustness; notably, stability under weak priors is markedly enhanced, and computational efficiency surpasses that of conventional ABC-RF.

Enables accurate parameter inference across diverse scientific domainsImproves ABC inference via random forests for complex modelsReduces computational cost and uncertainty in posterior estimation

Latest Papers

What's happening recently
View more

This study addresses the problem of ill-defined statistical semantics caused by zero-probability observations in probabilistic programs. Drawing on geometric measure theory, this work constructs a disintegration-based integral semantics over program traces. By formulating an explicit Bayesian conditioning rule compatible with mainstream inference algorithms, it rigorously rectifies the theoretical deficiencies inherent in existing soft conditioning approaches. The proposed framework naturally accommodates loops, mixture distributions, and manifold-valued observations. Consequently, it ensures both mathematical rigor and statistical correctness for conditional reasoning and inference in complex probabilistic models.

Bayes RuleConditioningProbabilistic Programs

This study addresses the high computational complexity of probabilistic inference and challenges in uncertainty modeling by proposing probabilistic circuits as a novel reasoning framework. By introducing structural constraints, the approach enables exact inference in polynomial time and innovatively integrates deep learning with symbolic paradigms to construct hybrid models supporting Bayesian learning. This research establishes a foundational theoretical system for probabilistic circuits, achieving efficient, exact computation and scalable deployment across diverse inference tasks. Ultimately, this work effectively bridges the gap between neural and symbolic AI, systematically advancing the development of probabilistic circuits at both theoretical and applied levels.

Probabilistic CircuitsProbabilistic InferenceTractability

This study addresses the unification of calibration concepts across classification and regression tasks, aiming to ensure consistency between predicted distributions and observed outcomes for diverse data types—continuous, discrete, nominal, and binary. The work introduces modal calibration for nominal outcomes and establishes a hierarchical framework distinguishing full, partial, and average calibration. It proposes a generalized definition of calibration based on predictive distribution functionals—such as means, quantiles, and event probabilities—and leverages probability integral transforms alongside constructive algorithms for analysis. Key contributions include demonstrating the logical independence between dual probability integral transform (PIT) calibration and existing discrete calibration notions, clarifying implication and independence relationships among various calibration types, and providing reproducible methods for generating illustrative examples and counterexamples.

calibrationclassificationhierarchical relations

This work addresses the issue of overconfidence in statistical inference arising from machine learning approximations in scientific simulations, which can compromise result reliability. To mitigate this, the authors propose two complementary approaches: first, a “balanced” regularization strategy that explicitly suppresses model overconfidence; and second, a simulation-aware Bayesian neural network prior that naturally alleviates overconfidence without additional regularization, even in small-sample regimes. By integrating neural ratio estimation with uncertainty quantification techniques, the proposed methods significantly improve inference calibration, yielding posterior estimates that are either closer to the ground truth or conservatively biased. This enhanced calibration strengthens the credibility of simulation-based inference in scientific applications.

calibrationmachine learningoverconfidence

This work addresses the NP-hard problem of learning polytree structures, systematically investigating the computational limits while preserving inference efficiency and interpretability. By incorporating indegree constraints, analyzing score function properties, and leveraging parameterized algorithms, combinatorial optimization, and approximation theory, the study presents an exact algorithm with time complexity $O((2+\varepsilon)^n)$, substantially improving upon the previous best-known bound of $O(3^n)$. Moreover, it introduces the first polynomial-time approximation algorithms: a $k$-approximation for general scoring functions and a tight 2-approximation for additive scores. Theoretical analysis demonstrates that most of these results are nearly optimal, establishing fundamental performance guarantees for polytree structure learning under practical constraints.

Bayesian networksin-degree boundsNP-hard

Hot Scholars

TV

Thibaut Vidal

Professor, SCALE-AI Chair, MAGI, Polytechnique Montréal
Combinatorial OptimizationMachine LearningOperations ResearchTransportation and Logistics
JR

Jesse Read

École Polytechnique
Multi-label ClassificationData-Stream LearningMachine LearningArtificial Intelligence
DB

Dimitris Bertsimas

Boeing Professor of Operations Research, MIT
Operations ResearchOptimizationStochasticsAnalytics
SM

Saumitra Mishra

Executive Director at JP Morgan
Explainable AIRobust Machine LearningMachine Learning Fairness
SI

Salim I. Amoukou

University Paris Saclay, J.P. Morgan
Uncertainty QuantificationInterpretable Machine LearningConformal PredictionDomain Adaptation