distributionally robust verification

Designs and analyzes verification procedures that compute rigorous worst-case bounds on the probability of specification violation by optimizing over all probability distributions within an ambiguity set consistent with available data or distributional constraints; the goal is to produce sound upper bounds on violation probability even under unknown or arbitrary correlations. This competency covers constructing ambiguity sets, formulating and solving the resulting distributionally robust optimization problems (including tractable relaxations), and implementing efficient algorithms suitable for tasks such as runtime monitoring and probabilistic guarantee computation.

distributionallyrobustverification

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Foundations for Deductive Verification of Continuous Probabilistic Programs: From Lebesgue to Riemann and Back

Feb 26, 2025
KB
Kevin Batz
🏛️ RWTH Aachen University | University College London | University of Trieste

This work addresses the automatic verification of expected output bounds for probabilistic programs featuring general loops, continuous distributions, and conditional branching—where the integral semantics induced by continuous sampling impede conventional invariant-based reasoning. We propose a Riemann-sum-based approximation of the expected semantics, transforming integral bounds into quantitative invariants expressible in SMT logic. This constitutes the first systematic integration of Riemann integration into probabilistic program verification, accompanied by formal convergence guarantees for the approximation and a proof that the verification problem is coRE-complete. We implement a prototype within the Caesar verification framework, supporting intermediate-language encoding and SMT-driven inference; it successfully verifies multiple benchmarks involving continuous sampling and loops. Our approach bridges discrete program verifiers with continuous probabilistic analysis, enabling existing discrete verification tools to scale to programs with continuous distributions.

Develops verification for continuous probabilistic programs.Enables SMT-based invariant verification.Uses Riemann sums to approximate integrals.

Ranking and Invariants for Lower-Bound Inference in Quantitative Verification of Probabilistic Programs

Apr 05, 2025
SK
Satoshi Kura
🏛️ Waseda University | Tohoku University | Chiba University

Inferring sound lower bounds for least fixed points in quantitative verification of probabilistic programs remains challenging. Method: We propose the first lower-bound verification framework grounded in uniqueness conditions, generalizing ranking functions from non-probabilistic programs to the probabilistic setting. Our approach establishes a theoretical connection between generalized ranking supermartingales and uniqueness of fixed points, ensuring correctness of inferred lower bounds. The framework unifies verification of diverse quantitative properties—including weak pre-expectations, expected runtime, and higher-order moments—by integrating template-based constraint solving with weakest preexpectation semantics. Results: We implement an automated verification tool based on this framework. Experimental evaluation demonstrates significant improvements in both effectiveness and precision of lower-bound inference across a broad suite of probabilistic programs, including those with complex control flow and stochastic dynamics.

Automating verification for quantitative properties using ranking supermartingalesEstimating lower bounds of least fixed points in probabilistic programsRelating fixed point uniqueness to program termination analysis

Quantitative Verification With Neural Networks For Probabilistic Programs and Stochastic Systems

Jan 15, 2023
AA
Alessandro Abate
🏛️ University of Oxford | University of Birmingham

This work addresses the quantitative verification of probabilistic programs and stochastic dynamical systems, specifically aiming to rigorously infer upper bounds on the probability that a stochastic process reaches a target condition within a finite number of steps. We propose a neuro-symbolic approach: supermartingale certificates are parameterized using differentiable neural networks; training employs stochastic optimization, while formal verification leverages SMT solvers (e.g., Z3); and an counterexample-guided inductive synthesis (CEGIS) framework enables iterative refinement. To our knowledge, this is the first method to embed neural networks directly into supermartingale construction—balancing expressive power with formal verifiability—and thereby significantly improves bound tightness and reliability. Evaluated on diverse benchmarks, our computed probability bounds match or surpass those of state-of-the-art techniques. Notably, we successfully verify high-dimensional, nonlinear stochastic models that defy analysis by conventional symbolic methods.

Computing tight probability bounds using neural networksHandling reachability, safety, and termination analysis efficientlyQuantitative verification of probabilistic programs and stochastic models

Sound Statistical Model Checking for Probabilities and Expected Rewards

Nov 01, 2024
CE
Carlos E. Budde
🏛️ University of Trento | University of Twente | Lancaster University Leipzig | Institute of Science and Technology Austria | Centre for Tactile Internet with Human-in-the-Loop (CeTI) | Technische Universität Dresden

Statistical Model Checking (SMC) often yields inflated error rates in probabilistic and expected reward estimation due to insufficient statistical rigor. To address this, we propose a robust estimation framework with rigorous theoretical guarantees: (i) we extend the Dvoretzky–Kiefer–Wolfowitz (DKW) inequality to expected reward estimation for the first time; (ii) we introduce a limit-PAC (Probably Approximately Correct) procedure ensuring controllable estimation error; and (iii) we derive a computable upper bound on reachability rewards and enhance practicality via path truncation and distribution bounding. Our method is implemented in the *modes* tool. Experimental evaluation demonstrates a substantial reduction in erroneous conclusions while maintaining high precision, thereby ensuring both statistical correctness and engineering applicability.

Developing sound methods for probability estimationEstablishing bounds for expected reward distributionsEvaluating correctness of statistical model checking tools

Latest Papers

What's happening recently
View more

This work addresses the challenge of silent errors in large language models (LLMs) when translating natural language into optimization models, which often lead to subtle semantic deviations. The authors propose a reference-free, falsification-based verification framework that generates test instances through slot perturbations and leverages solvers for error detection. Drawing on duality theory, comparative statics, and polyhedral limit arguments, they design a suite of acoustic test classes—including collapse probes, forbidden limits, annihilation, and swap tests—that jointly examine directional, curvature, and asymptotic properties. This approach achieves, for the first time, zero false positives in verifying model fidelity. Evaluated on 326 real-world models, the method attains a 0.0% false positive rate—substantially outperforming threshold-based methods (54.9%)—and detects 70.0% of conditional mutants, 40.4% of which are invisible to conventional execution-based accuracy metrics.

falsificationlarge language modelsoptimization models

This study addresses the longstanding challenge of lacking formal connections between search and refutation algorithms for random constraint satisfaction problems (CSPs). By leveraging average-case complexity theory and semi-random model analysis, this work constructs a unified theoretical framework that rigorously relates search and refutation processes for the first time. It demonstrates that existing algorithms can output near-optimal solutions accompanied by certificates, and proposes optimization methods with verifiable ε-optimality guarantees. Furthermore, novel robust algorithms are designed to operate under strong contamination models. Ultimately, this research achieves verifiably near-optimal solving for semi-random and contaminated CSPs, effectively unifying computational threshold theory and providing a new paradigm that combines theoretical rigor with practical utility for related fields.

Certifiable Near-OptimalityRandom CSPsSearch and Refutation

This work addresses the challenge of verifying probabilistic safety policies in AI agents, which existing runtime monitoring approaches struggle to handle due to their reliance on deterministic strategies and inability to account for uncertainty and correlations. The paper proposes a novel verification framework based on distributionally robust optimization (DRO), introducing this technique for the first time into Datalog-style probabilistic policy verification. By eschewing assumptions of predicate independence, the method provides rigorous upper bounds on policy violation probabilities. Integrating DRO with probabilistic logical reasoning and Datalog’s formalism, the approach achieves a superior trade-off between safety and utility while maintaining strict guarantees on violation probability bounds. Empirical evaluations on benchmarks involving terminal and tool-calling agents demonstrate significant improvements over state-of-the-art methods.

AI agentsDatalogdistributional robustness

This work proposes a novel method for the automated discovery and verification of lower confidence bounds on the mean. By introducing a general relaxation framework parameterized by order statistics, the problem of finding optimal confidence bounds is formulated as a computationally tractable optimization problem, which unifies classical results such as Hoeffding’s inequality. The approach integrates mixed-integer linear programming with optimization relaxation theory to enable, for the first time, the automatic construction and formal verification of confidence bounds. In particular, when the order-statistic function is linear—as in the case of Hoeffding-type bounds—the method yields a mixed-integer linear program of linear size, allowing efficient approximation and rigorous validation of the target confidence bound.

automated provingconfidence boundsmean estimation

This study addresses the challenge that existing methods struggle to provide tight and verifiable generalization error bounds for modern algorithms. We propose a theoretical framework grounded in formal methods that recasts algorithmic stability as system specifications and computes provable generalization bounds via reachability analysis. Furthermore, we construct novel concentration inequalities to bridge sample-level results with distribution-level analysis, enabling rigorous verification of boundedness without requiring analytic assumptions. By integrating formal verification, reachability analysis, and algorithmic stability theory, our approach yields tighter and computable generalization guarantees compared to traditional methods, particularly in complex scenarios such as large language model fine-tuning.

algorithmic stabilityconcentration inequalityexpected generalization gap

Hot Scholars

JP

Joost-Pieter Katoen

Distinguished Professor of Computer Science, RWTH Aachen University and University of Twente
formal methodsmodel checkingconcurrency theoryprobabilistic programming
TQ

Tim Quatmann

RWTH Aachen University
Artificial IntelligenceFormal MethodsModel Checking
SD

Steven Dillmann

Stanford University, University of Cambridge
AI for ScienceMachine LearningData Driven DiscoveryComputational Mathematics
JZ

Jiacheng Zhu

MIT
Machine LearningFoundation ModelsOptimal TransportBayesian modeling
KK

Karl Krauth

Postdoc, Stanford
machine learningstatisticsoptimization