Score
Designs and analyzes procedures that guarantee, with high probability in finite samples, that prediction accuracy or constraint violations meet specified targets; this includes constructing finite-sample confidence bounds, calibrating decision thresholds, and producing probabilistic guarantees for constraint satisfaction.
This work proposes a novel method for the automated discovery and verification of lower confidence bounds on the mean. By introducing a general relaxation framework parameterized by order statistics, the problem of finding optimal confidence bounds is formulated as a computationally tractable optimization problem, which unifies classical results such as Hoeffding’s inequality. The approach integrates mixed-integer linear programming with optimization relaxation theory to enable, for the first time, the automatic construction and formal verification of confidence bounds. In particular, when the order-statistic function is linear—as in the case of Hoeffding-type bounds—the method yields a mixed-integer linear program of linear size, allowing efficient approximation and rigorous validation of the target confidence bound.
Statistical Model Checking (SMC) often yields inflated error rates in probabilistic and expected reward estimation due to insufficient statistical rigor. To address this, we propose a robust estimation framework with rigorous theoretical guarantees: (i) we extend the Dvoretzky–Kiefer–Wolfowitz (DKW) inequality to expected reward estimation for the first time; (ii) we introduce a limit-PAC (Probably Approximately Correct) procedure ensuring controllable estimation error; and (iii) we derive a computable upper bound on reachability rewards and enhance practicality via path truncation and distribution bounding. Our method is implemented in the *modes* tool. Experimental evaluation demonstrates a substantial reduction in erroneous conclusions while maintaining high precision, thereby ensuring both statistical correctness and engineering applicability.
For Markov decision processes (MDPs) with unknown transition probabilities, existing statistical model checking (SMC) algorithms suffer from high sample complexity and weak theoretical guarantees. Method: We introduce tight concentration inequalities—specifically, the Bretagnolle–Huber and Empirical Bernstein bounds—into the SMC framework for the first time, and design adaptive, structure-aware statistical estimators that exploit MDP topology. Contribution/Results: Theoretically, our approach yields significantly tighter and more general probably approximately correct (PAC) guarantees. Empirically, it reduces required sample sizes by up to two orders of magnitude on standard verification benchmarks. This work establishes a new paradigm for efficient and reliable formal verification of uncertain systems.
This work addresses the challenge of quantifying prediction uncertainty in generative biomolecular design, where feedback covariate shift undermines conventional uncertainty estimation. We propose the first conformal prediction framework tailored to closed-loop design paradigms. Departing from standard i.i.d. assumptions, our method imposes no structural constraints on either the design algorithm or the regression model, delivering finite-sample statistically valid confidence sets for arbitrary black-box design pipelines. Key innovations include quantile-regression-driven adaptive conformal prediction, explicit modeling of feedback-induced distributional shift, and robust error calibration. Evaluated on protein and small-molecule design tasks, our approach achieves ≥94.8% empirical coverage at the 95% nominal confidence level—substantially outperforming standard conformal methods (which drop to as low as 72%)—while maintaining high predictive accuracy.
This work addresses the quantitative verification of probabilistic programs and stochastic dynamical systems, specifically aiming to rigorously infer upper bounds on the probability that a stochastic process reaches a target condition within a finite number of steps. We propose a neuro-symbolic approach: supermartingale certificates are parameterized using differentiable neural networks; training employs stochastic optimization, while formal verification leverages SMT solvers (e.g., Z3); and an counterexample-guided inductive synthesis (CEGIS) framework enables iterative refinement. To our knowledge, this is the first method to embed neural networks directly into supermartingale construction—balancing expressive power with formal verifiability—and thereby significantly improves bound tightness and reliability. Evaluated on diverse benchmarks, our computed probability bounds match or surpass those of state-of-the-art techniques. Notably, we successfully verify high-dimensional, nonlinear stochastic models that defy analysis by conventional symbolic methods.
This study addresses the longstanding challenge of lacking formal connections between search and refutation algorithms for random constraint satisfaction problems (CSPs). By leveraging average-case complexity theory and semi-random model analysis, this work constructs a unified theoretical framework that rigorously relates search and refutation processes for the first time. It demonstrates that existing algorithms can output near-optimal solutions accompanied by certificates, and proposes optimization methods with verifiable ε-optimality guarantees. Furthermore, novel robust algorithms are designed to operate under strong contamination models. Ultimately, this research achieves verifiably near-optimal solving for semi-random and contaminated CSPs, effectively unifying computational threshold theory and providing a new paradigm that combines theoretical rigor with practical utility for related fields.
本文提出一种基于置信区域的筛选框架,用于解决模拟系统可接受性问题,保证高概率筛选出所有或每个可接受系统,并支持并行化。
This work addresses the problem of efficiently verifying the consistency of a large number of conditional probability assertions generated by probabilistic predictors. To this end, we construct an interactive PCP protocol in which a polynomial-time verifier need only make a small number of queries to a probabilistic circuit (P, Q) and read a few positions of a proof oracle to check its ℓ²-approximate consistency. Our approach is the first to place the decision problem for explicit probabilistic assertion consistency within NP, eliminating dependence on input precision. By leveraging sparse witness distributions and circuit-based probabilistic modeling, this work establishes a rigorous complexity-theoretic foundation for self-consistency verification in predictive models, enabling efficient validation of exponentially many implicit assertions.
This work addresses the lack of statistical reliability guarantees in existing hyperparameter selection methods—such as grid search and Bayesian optimization—with respect to critical metrics like risk and safety. Building upon the learn-then-test (LTT) paradigm, the paper introduces a unified statistical framework that formulates hyperparameter selection as a multiple hypothesis testing problem, accommodating user-specified constraints on average risk, quantile risk, or information-theoretic measures. Leveraging tools from statistical inference—including p-values, e-values, and concentration inequalities—the method derives explicit, finite-sample bounds on error probabilities from first principles. This approach enables theoretically grounded validation and selection of hyperparameters, substantially enhancing the reliability and safety of AI systems in real-world deployment scenarios.
This study addresses the challenge of efficiently comparing and selecting among multiple adaptive prediction pipelines under hard coverage constraints. To this end, it proposes the CC-SMCS framework, which achieves sequential model selection by decoupling feasibility from optimality. Technically, the approach employs stochastic constrained argmin modeling, simultaneous martingale confidence sequences, and rectangular region projections, yielding closed-form rules with finite-sample guarantees without requiring stationarity assumptions. The proposed method contains all constrained optimal pipelines with high probability while supporting data-dependent stopping and delayed feedback scenarios. Furthermore, this work establishes an impossibility result for margin-safe certification, thereby providing a theoretical foundation for constrained online learning.