Score
Deriving and proving formal statistical properties (identifiability, asymptotic guarantees, screening-safety) and the conditions on data, models, or priors under which estimators and recommended treatments are valid.
This study addresses the problem of verifying whether a given observational formula correctly identifies a target interventional distribution in causal graphical models, going beyond mere identifiability assessment. To this end, it introduces a falsification-driven verification framework that decouples verification from identification for the first time: an efficient falsifier first eliminates incorrect formulas, and a verifier—provably almost surely correct under regular exponential family models—is then constructed atop this filter. As an application, the authors develop a “gateway test” that enumerates all valid variable sets satisfying the front-door criterion, with theoretical guarantees on verification reliability. This approach substantially enhances both the practicality and rigor of validating interventional distributions.
This study investigates the trade-off degradation between parameter identifiability and predictive falsifiability in Bayesian model extensions. We formally establish, for the first time, their intrinsic negative correlation—where increased model complexity simultaneously degrades both properties. To mitigate this tension, we propose a novel inference framework grounded in the posterior joint structure of parameters and predictions. Through theoretical analysis and two canonical extension examples, we demonstrate that our approach improves the synergistic balance: it enhances the uniqueness of parameter interpretation (identifiability) while preserving empirical testability of predictions (falsifiability). Our core contribution is the introduction of a unified identifiability–falsifiability diagnostic perspective, providing a new paradigm for Bayesian modeling that integrates statistical rigor with scientific testability.
This work addresses the formal verification of Markov processes parameterized by machine learning models—including linear regressors, decision trees, and neural networks—to ensure reliability in safety-critical applications such as medical modeling and probabilistic programming. Methodologically, we embed ML parameters into Markov processes and formulate a bilinear programming model; we further propose a novel parameter decomposition technique coupled with interval-bound propagation to enable efficient, globally optimal verification of key properties—including reachability, hitting time, and total reward. Our approach achieves up to 100× speedup over state-of-the-art solvers. We release MarkovML, an open-source tool supporting high-level modeling, seamless ML integration, and end-to-end automated verification. This framework advances formal analysis of AI-augmented stochastic systems, bridging the gap between learning-based components and rigorous probabilistic guarantees.
Hypothesis testing in singular statistical models is often deemed infeasible due to non-identifiable parameters and degenerate Fisher information. This work circumvents these issues by reframing hypotheses in terms of identifiable functionals of the observable distribution rather than unidentifiable parameter functions, thereby recasting the problem as a classical testing problem in the space of probability distributions. The study introduces the novel concept of an “overlap barrier,” which reveals that hypotheses involving non-identifiable quantities inevitably lead to testing impossibility. A Hellinger-distance-based criterion for testability is established, enabling a structural classification of hypotheses in singular models. By integrating distribution-space analysis, posterior contraction theory, and test-driven Bayesian arguments, the framework is validated in Gaussian mixture models and reduced-rank regression, rigorously delineating the boundary between testable and non-testable hypotheses and clarifying the limits of valid statistical inference in singular settings.
This work addresses the structural identifiability of parameters in stochastic differential equation (SDE) models under multiple interventions—i.e., whether SDE parameters can be uniquely recovered from samples of post-intervention stationary distributions. Theoretically, we establish the first uniqueness guarantee for SDE parameter recovery under multi-intervention settings; for linear SDEs, we derive a tight lower bound on the minimum number of required interventions; for weak-noise nonlinear SDEs, we obtain an upper bound on identifiability. Methodologically, we propose a parametric framework featuring learnable activation functions, integrating intervention modeling, stationary distribution analysis, and weak-noise asymptotic theory. Experiments on synthetic data demonstrate that our approach accurately recovers ground-truth parameters, and the theory-guided learnable architecture significantly improves both estimation accuracy and robustness.
This study addresses the fundamental question of whether a statistical parameter defined through a conditional distribution remains constant across covariates—a problem encompassing treatment effect heterogeneity and conditional association. The authors propose a general nonparametric testing framework based on smooth functionals applied to conditional distributions, yielding functional parameters for which they construct test statistics with tractable asymptotic distributions. Their approach explicitly links to norm-based tests in function spaces and, compared to existing norm-type methods, exhibits superior asymptotic properties under the null hypothesis. Simulation studies demonstrate strong finite-sample performance, and the method is successfully applied to data from a breast cancer clinical trial, effectively identifying key biomarkers predictive of response to adjuvant chemotherapy.
Traditional statistical model checking (SMC) is limited to mean-type statistics—such as expected reward and probability—and thus struggles to reliably estimate distribution-sensitive risk measures, including quantiles, conditional value-at-risk (CVaR), and entropy-based risk. This work introduces, for the first time, the Dvoretzky–Kiefer–Wolfowitz–Massart (DKW) inequality into SMC, enabling non-asymptotic confidence bands around the empirical cumulative distribution function (ECDF) to perform rigorous, distribution-wide inference over system trajectories. The resulting framework supports statistical verification of key risk metrics—quantiles, CVaR, and entropy risk—with formal guarantees. It is efficiently implemented in the MODEST Toolset’s modes simulator. Experimental evaluation across multiple quantitative verification benchmarks demonstrates high accuracy and strong robustness, significantly extending the risk modeling capabilities and expressive power of SMC for reliability assessment.
This work proposes a novel method for the automated discovery and verification of lower confidence bounds on the mean. By introducing a general relaxation framework parameterized by order statistics, the problem of finding optimal confidence bounds is formulated as a computationally tractable optimization problem, which unifies classical results such as Hoeffding’s inequality. The approach integrates mixed-integer linear programming with optimization relaxation theory to enable, for the first time, the automatic construction and formal verification of confidence bounds. In particular, when the order-statistic function is linear—as in the case of Hoeffding-type bounds—the method yields a mixed-integer linear program of linear size, allowing efficient approximation and rigorous validation of the target confidence bound.
This study addresses the joint identification and counterfactual analysis in incomplete structural models featuring support and moment constraints. The authors embed counterfactuals directly into an augmented structural model, departing from the conventional “estimate-then-simulate” paradigm. By leveraging support function methods, they simultaneously achieve identification and inference, revealing a fundamental isomorphism between the two tasks. A key contribution is the formulation of irreducibility conditions that explicitly characterize all support implications. Under mild regularity assumptions, the support function approach preserves sharpness with respect to the moment closure—even in counterfactual settings where traditional sharpness fails. Moreover, for irreducible models, the identified set and the moment closure are statistically indistinguishable in finite samples.
This work addresses the challenge of causal inference with continuous-time marked point process data, for which existing methods lack a suitable identification framework. Building on martingale theory, the authors extend the core assumptions of discrete-time causal inference—consistency, exchangeability, and positivity—to the continuous-time setting. They formulate a dynamic treatment strategy and a potential outcomes model tailored to marked point processes and establish corresponding causal identification conditions. Leveraging this foundation, they derive a novel marginal g-formula that enables nonparametric identification of causal effects. The proposed framework subsumes existing results for discrete-time and counting process settings as special cases, demonstrating both theoretical compatibility and extensibility, thereby unifying survival analysis and causal inference within a coherent paradigm.