Score
Designs and analyzes verification procedures that compute rigorous worst-case bounds on the probability of specification violation by optimizing over all probability distributions within an ambiguity set consistent with available data or distributional constraints; the goal is to produce sound upper bounds on violation probability even under unknown or arbitrary correlations. This competency covers constructing ambiguity sets, formulating and solving the resulting distributionally robust optimization problems (including tractable relaxations), and implementing efficient algorithms suitable for tasks such as runtime monitoring and probabilistic guarantee computation.
This work addresses the automatic verification of expected output bounds for probabilistic programs featuring general loops, continuous distributions, and conditional branching—where the integral semantics induced by continuous sampling impede conventional invariant-based reasoning. We propose a Riemann-sum-based approximation of the expected semantics, transforming integral bounds into quantitative invariants expressible in SMT logic. This constitutes the first systematic integration of Riemann integration into probabilistic program verification, accompanied by formal convergence guarantees for the approximation and a proof that the verification problem is coRE-complete. We implement a prototype within the Caesar verification framework, supporting intermediate-language encoding and SMT-driven inference; it successfully verifies multiple benchmarks involving continuous sampling and loops. Our approach bridges discrete program verifiers with continuous probabilistic analysis, enabling existing discrete verification tools to scale to programs with continuous distributions.
Inferring sound lower bounds for least fixed points in quantitative verification of probabilistic programs remains challenging. Method: We propose the first lower-bound verification framework grounded in uniqueness conditions, generalizing ranking functions from non-probabilistic programs to the probabilistic setting. Our approach establishes a theoretical connection between generalized ranking supermartingales and uniqueness of fixed points, ensuring correctness of inferred lower bounds. The framework unifies verification of diverse quantitative properties—including weak pre-expectations, expected runtime, and higher-order moments—by integrating template-based constraint solving with weakest preexpectation semantics. Results: We implement an automated verification tool based on this framework. Experimental evaluation demonstrates significant improvements in both effectiveness and precision of lower-bound inference across a broad suite of probabilistic programs, including those with complex control flow and stochastic dynamics.
本文解决了电路最小化中无法验证最优性声明的问题,通过合成GF(2)上的最小线性程序并生成可验证的DRAT证明来确保声明的两部分都得到认证。
This work addresses the quantitative verification of probabilistic programs and stochastic dynamical systems, specifically aiming to rigorously infer upper bounds on the probability that a stochastic process reaches a target condition within a finite number of steps. We propose a neuro-symbolic approach: supermartingale certificates are parameterized using differentiable neural networks; training employs stochastic optimization, while formal verification leverages SMT solvers (e.g., Z3); and an counterexample-guided inductive synthesis (CEGIS) framework enables iterative refinement. To our knowledge, this is the first method to embed neural networks directly into supermartingale construction—balancing expressive power with formal verifiability—and thereby significantly improves bound tightness and reliability. Evaluated on diverse benchmarks, our computed probability bounds match or surpass those of state-of-the-art techniques. Notably, we successfully verify high-dimensional, nonlinear stochastic models that defy analysis by conventional symbolic methods.
Statistical Model Checking (SMC) often yields inflated error rates in probabilistic and expected reward estimation due to insufficient statistical rigor. To address this, we propose a robust estimation framework with rigorous theoretical guarantees: (i) we extend the Dvoretzky–Kiefer–Wolfowitz (DKW) inequality to expected reward estimation for the first time; (ii) we introduce a limit-PAC (Probably Approximately Correct) procedure ensuring controllable estimation error; and (iii) we derive a computable upper bound on reachability rewards and enhance practicality via path truncation and distribution bounding. Our method is implemented in the *modes* tool. Experimental evaluation demonstrates a substantial reduction in erroneous conclusions while maintaining high precision, thereby ensuring both statistical correctness and engineering applicability.
This work addresses the challenge of silent errors in large language models (LLMs) when translating natural language into optimization models, which often lead to subtle semantic deviations. The authors propose a reference-free, falsification-based verification framework that generates test instances through slot perturbations and leverages solvers for error detection. Drawing on duality theory, comparative statics, and polyhedral limit arguments, they design a suite of acoustic test classes—including collapse probes, forbidden limits, annihilation, and swap tests—that jointly examine directional, curvature, and asymptotic properties. This approach achieves, for the first time, zero false positives in verifying model fidelity. Evaluated on 326 real-world models, the method attains a 0.0% false positive rate—substantially outperforming threshold-based methods (54.9%)—and detects 70.0% of conditional mutants, 40.4% of which are invisible to conventional execution-based accuracy metrics.
This study addresses the longstanding challenge of lacking formal connections between search and refutation algorithms for random constraint satisfaction problems (CSPs). By leveraging average-case complexity theory and semi-random model analysis, this work constructs a unified theoretical framework that rigorously relates search and refutation processes for the first time. It demonstrates that existing algorithms can output near-optimal solutions accompanied by certificates, and proposes optimization methods with verifiable ε-optimality guarantees. Furthermore, novel robust algorithms are designed to operate under strong contamination models. Ultimately, this research achieves verifiably near-optimal solving for semi-random and contaminated CSPs, effectively unifying computational threshold theory and providing a new paradigm that combines theoretical rigor with practical utility for related fields.
This work addresses the challenge of verifying probabilistic safety policies in AI agents, which existing runtime monitoring approaches struggle to handle due to their reliance on deterministic strategies and inability to account for uncertainty and correlations. The paper proposes a novel verification framework based on distributionally robust optimization (DRO), introducing this technique for the first time into Datalog-style probabilistic policy verification. By eschewing assumptions of predicate independence, the method provides rigorous upper bounds on policy violation probabilities. Integrating DRO with probabilistic logical reasoning and Datalog’s formalism, the approach achieves a superior trade-off between safety and utility while maintaining strict guarantees on violation probability bounds. Empirical evaluations on benchmarks involving terminal and tool-calling agents demonstrate significant improvements over state-of-the-art methods.
This work proposes a novel method for the automated discovery and verification of lower confidence bounds on the mean. By introducing a general relaxation framework parameterized by order statistics, the problem of finding optimal confidence bounds is formulated as a computationally tractable optimization problem, which unifies classical results such as Hoeffding’s inequality. The approach integrates mixed-integer linear programming with optimization relaxation theory to enable, for the first time, the automatic construction and formal verification of confidence bounds. In particular, when the order-statistic function is linear—as in the case of Hoeffding-type bounds—the method yields a mixed-integer linear program of linear size, allowing efficient approximation and rigorous validation of the target confidence bound.
This study addresses the challenge that existing methods struggle to provide tight and verifiable generalization error bounds for modern algorithms. We propose a theoretical framework grounded in formal methods that recasts algorithmic stability as system specifications and computes provable generalization bounds via reachability analysis. Furthermore, we construct novel concentration inequalities to bridge sample-level results with distribution-level analysis, enabling rigorous verification of boundedness without requiring analytic assumptions. By integrating formal verification, reachability analysis, and algorithmic stability theory, our approach yields tighter and computable generalization guarantees compared to traditional methods, particularly in complex scenarios such as large language model fine-tuning.