Score
Designs, builds, or analyzes decision-tree models and inference procedures that represent and propagate probability distributions through nodes and soft branchings to produce calibrated probabilistic predictions and uncertainty estimates while preserving interpretable decision paths. Work includes specifying probabilistic routing or soft-split rules, implementing algorithms to compute posterior or predictive distributions through the tree, and evaluating calibration and interpretability of the resulting probabilistic tree models.
To address the growing demand for trustworthy AI, this work proposes a calibratable nonparametric probabilistic regression tree that directly models the conditional cumulative distribution function (cCDF) to produce prediction intervals with high accuracy, reliable coverage, and statistical calibration. Methodologically, it introduces efficient tree-splitting algorithms tailored to the weighted interval score (WIS) and continuous ranked probability score (CRPS)—the first such formulations for these proper scoring rules. To accelerate split search and gradient updates, it synergistically integrates min-max heaps, weight-balanced binary trees, and Fenwick trees. Additionally, the framework supports group-conditional conformal prediction, ensuring statistical validity while preserving model interpretability. Empirical evaluation demonstrates superior calibration, coverage, and computational efficiency compared to state-of-the-art baselines across diverse benchmarks. The method achieves a balanced advantage in critical scientific and engineering applications where both uncertainty quantification rigor and practical deployability are essential.
本文通过将概率视为预测方法的输出,统一了不同概率解释,并提出当满足有限校准标准时,可以预测效用分布以指导决策。
In same-domain repeated access scenarios, decision trees and their ensembles (e.g., random forests, GBDTs) struggle to jointly optimize accuracy and interpretability. Method: This paper proposes a learnable, tunable unified tree framework. It introduces (1) a parameterized splitting criterion that continuously interpolates between entropy and Gini impurity, enabling data-adaptive optimal splits; (2) a theoretical characterization of sample complexity, providing generalization guarantees for the interpretability–accuracy trade-off; and (3) joint optimization of Bayesian decision trees, minimum-cost-complexity pruning, and ensemble hyperparameters. Results: Extensive experiments on real-world datasets demonstrate that the framework significantly improves the consistency between predictive accuracy and model interpretability. It offers both theoretical rigor—via provable generalization bounds—and practical utility—through end-to-end differentiability and seamless integration into existing tree-based pipelines. The approach bridges a critical gap between statistical performance and human-understandable structure in tree learning.
For Markov decision processes (MDPs) with unknown transition probabilities, existing statistical model checking (SMC) algorithms suffer from high sample complexity and weak theoretical guarantees. Method: We introduce tight concentration inequalities—specifically, the Bretagnolle–Huber and Empirical Bernstein bounds—into the SMC framework for the first time, and design adaptive, structure-aware statistical estimators that exploit MDP topology. Contribution/Results: Theoretically, our approach yields significantly tighter and more general probably approximately correct (PAC) guarantees. Empirically, it reduces required sample sizes by up to two orders of magnitude on standard verification benchmarks. This work establishes a new paradigm for efficient and reliable formal verification of uncertain systems.
In likelihood-intractable scenarios, Approximate Bayesian Computation (ABC) faces key bottlenecks: difficulty in selecting informative summary statistics, complex tuning of distance metrics and tolerance thresholds, high computational cost, and substantial posterior uncertainty due to weak prior information. To address these challenges, this paper proposes ABC-SMC-(D)RF—a novel method integrating Distributed Random Forests (DRF) into the Sequential Monte Carlo (SMC) framework for the first time. It enables end-to-end learning of discriminative features, eliminating manual design of summary statistics and distance functions. Coupled with adaptive tolerance scheduling and a pre-acceptance shrinkage mechanism, it iteratively concentrates sampling on high-probability parameter regions. Experiments across diverse deterministic and stochastic models demonstrate significant improvements in posterior estimation accuracy and robustness; notably, stability under weak priors is markedly enhanced, and computational efficiency surpasses that of conventional ABC-RF.
This study addresses the problem of ill-defined statistical semantics caused by zero-probability observations in probabilistic programs. Drawing on geometric measure theory, this work constructs a disintegration-based integral semantics over program traces. By formulating an explicit Bayesian conditioning rule compatible with mainstream inference algorithms, it rigorously rectifies the theoretical deficiencies inherent in existing soft conditioning approaches. The proposed framework naturally accommodates loops, mixture distributions, and manifold-valued observations. Consequently, it ensures both mathematical rigor and statistical correctness for conditional reasoning and inference in complex probabilistic models.
This study addresses the high computational complexity of probabilistic inference and challenges in uncertainty modeling by proposing probabilistic circuits as a novel reasoning framework. By introducing structural constraints, the approach enables exact inference in polynomial time and innovatively integrates deep learning with symbolic paradigms to construct hybrid models supporting Bayesian learning. This research establishes a foundational theoretical system for probabilistic circuits, achieving efficient, exact computation and scalable deployment across diverse inference tasks. Ultimately, this work effectively bridges the gap between neural and symbolic AI, systematically advancing the development of probabilistic circuits at both theoretical and applied levels.
This study addresses the unification of calibration concepts across classification and regression tasks, aiming to ensure consistency between predicted distributions and observed outcomes for diverse data types—continuous, discrete, nominal, and binary. The work introduces modal calibration for nominal outcomes and establishes a hierarchical framework distinguishing full, partial, and average calibration. It proposes a generalized definition of calibration based on predictive distribution functionals—such as means, quantiles, and event probabilities—and leverages probability integral transforms alongside constructive algorithms for analysis. Key contributions include demonstrating the logical independence between dual probability integral transform (PIT) calibration and existing discrete calibration notions, clarifying implication and independence relationships among various calibration types, and providing reproducible methods for generating illustrative examples and counterexamples.
This work addresses the issue of overconfidence in statistical inference arising from machine learning approximations in scientific simulations, which can compromise result reliability. To mitigate this, the authors propose two complementary approaches: first, a “balanced” regularization strategy that explicitly suppresses model overconfidence; and second, a simulation-aware Bayesian neural network prior that naturally alleviates overconfidence without additional regularization, even in small-sample regimes. By integrating neural ratio estimation with uncertainty quantification techniques, the proposed methods significantly improve inference calibration, yielding posterior estimates that are either closer to the ground truth or conservatively biased. This enhanced calibration strengthens the credibility of simulation-based inference in scientific applications.
This work addresses the NP-hard problem of learning polytree structures, systematically investigating the computational limits while preserving inference efficiency and interpretability. By incorporating indegree constraints, analyzing score function properties, and leveraging parameterized algorithms, combinatorial optimization, and approximation theory, the study presents an exact algorithm with time complexity $O((2+\varepsilon)^n)$, substantially improving upon the previous best-known bound of $O(3^n)$. Moreover, it introduces the first polynomial-time approximation algorithms: a $k$-approximation for general scoring functions and a tight 2-approximation for additive scores. Theoretical analysis demonstrates that most of these results are nearly optimal, establishing fundamental performance guarantees for polytree structure learning under practical constraints.