Score
Designs, implements, or analyzes randomized-smoothing certification procedures that remain statistically valid at arbitrary stopping times by using sequential sampling and e-process (anytime-valid) tests; this includes constructing robustness certificates, concentration bounds, and decision rules that permit early stopping and adaptive allocation of compute per input while preserving rigorous validity guarantees.
This work addresses the problem of real-time, data-adaptive lower bounding of the number of true discoveries in online multiple testing, under the constraint that decisions must be based solely on past hypotheses and data—without lookahead. To tackle this challenge, we first establish that admissibility of online closed testing is equivalent to the availability of valid e-values at each time step. Building on this equivalence, we propose a novel online closed testing framework grounded in products of e-values, accompanied by an efficient algorithm and new e-value hedging and boosting mechanisms to enhance statistical power. Our method unifies modeling of exchangeable and arbitrarily dependent test statistics, ensuring strict control of both false discovery rate (FDR) and false discovery exceedance (FDX) under arbitrary dependence, while delivering tight lower confidence bounds on the number of true discoveries. Compared to state-of-the-art approaches, our method achieves theoretical unification and improved performance, enabling the first online procedure for true discovery control that simultaneously delivers strong empirical power and rigorous finite-sample guarantees.
Traditional fixed-sample hypothesis tests lack temporal flexibility—they cannot be terminated early or extended adaptively without inflating Type I error. Method: We propose an anytime-valid sequential testing framework that preserves statistical power. We rigorously prove that any fixed-sample test can be equivalently transformed into a sequential counterpart with identical power. This is achieved via p-value reconstruction and reinterpretation of significance levels, ensuring that the test can be stopped or continued at any time while strictly controlling the overall Type I error rate. Contributions/Results: We derive explicit anytime-valid versions of the z-test and t-test, which coincide exactly with their classical fixed-sample counterparts after N observations. We further show that the log-optimal sequential z-test corresponds to rejecting the null at the minimal future significance level required by the standard z-test. Our framework unifies fixed-sample and sequential paradigms, enabling reliable inference under dynamic, real-time data collection.
In sequential anytime testing, combining e-process evidence across filtrations poses a fundamental challenge: an e-process valid under a coarse filtration may fail to retain validity under a finer one, even for the same null hypothesis. Method: We propose the “adjusted combination” framework, introducing the class of *adjuster* functions—characterizing necessary and sufficient conditions for cross-filtration evidence boosting. We prove that an adjuster is essential to restore anytime validity under the original filtration and quantify its logarithmic cost. Our approach unifies e-process theory, generalized test martingales, and filtration coarsening/refinement techniques. Results: We establish a complete characterization theorem for adjusters and validate the framework on real financial data for randomness testing. The work provides novel theoretical tools and practical methodology for sequential independence testing and predictive model evaluation.
This work addresses the challenge of certifying robustness against multi-step, input-dependent noise perturbations under test-time adaptive defenses—a setting where conventional randomized smoothing fails due to its inability to handle structured, adaptive disturbances. We propose the first adaptive randomized smoothing framework grounded in *f*-differential privacy, enabling provably sound compositional certification for high-dimensional input-dependent masking and multi-step dynamic noise injection. Our method integrates *f*-DP analysis, adaptive noise modeling, and a novel multi-step smoothing certification mechanism. Evaluated on CIFAR-10 and CelebA, it improves standard accuracy by 1–15 percentage points; on ImageNet, certified accuracy increases by up to 1.6 percentage points. The framework significantly enhances both adversarial robustness and practical deployability under adaptive defenses.
This work proposes a flexible and rigorous monitoring framework for two-arm randomized controlled trials that addresses the challenge of Type I error control in adaptive designs with frequent interim analyses and data-dependent adaptations. Built upon E-values and E-processes, the approach enables valid inference under composite null hypotheses and supports futility monitoring, seamlessly integrating group sequential and Bayesian perspectives. By constructing E-processes via betting martingales and incorporating calibration strategies, multiplicity adjustments, and hybrid design elements, the method guarantees strict Type I error control without requiring pre-specified analysis times. The framework is implemented in the open-source R package evalinger. Numerical experiments demonstrate that, under continuous monitoring, the proposed method not only maintains exact Type I error control but also achieves higher statistical power compared to conventional group sequential approaches.
This work addresses the high computational cost and inflexibility of randomized smoothing (RS), which, despite offering rigorous robustness guarantees, requires a preset number of samples and is ill-suited for real-time or resource-constrained settings. The paper introduces anytime-valid robustness certification—a novel paradigm enabled by a meta-learning-based adaptive framework. A lightweight meta-learner predicts input-specific priors to guide a sequential estimation process that dynamically allocates sampling resources and supports early stopping. This approach preserves statistical validity while drastically reducing sampling complexity—by up to 20× compared to conventional RS—and enables risk-aware, tiered allocation of computation based on user-defined confidence thresholds. The method thus opens a new pathway toward efficient, real-time robustness certification in safety-critical applications.
This study addresses the invalidation of Gaussian process (GP) confidence envelope certificates caused by adaptive hyperparameter tuning in black-box optimization. To overcome this, the work introduces sequential e-values into GP model selection for the first time, dynamically eliminating contradictory candidates within a reproducing kernel Hilbert space to construct an auditable GP selection mechanism. The proposed approach provides anytime-valid near-optimality certification, resolving the challenge of maintaining statistical validity under adaptive evaluations. Empirical results demonstrate that, in noisy testing environments, the false positive rate is halved compared to fixed pre-commitment strategies, significantly enhancing both certification efficiency and robustness.
This study addresses the challenge of false positive accumulation caused by verifiers in adaptive agent generation, as well as the absence of reliable stopping criteria within generate-verify loops. To overcome these limitations, this work proposes an e-value analytical framework based on exponential betting, coupled with a novel conformal risk control procedure tailored for non-monotonic losses. By integrating distribution-free statistical testing, the proposed theory rigorously bounds the false discovery rate (FDR) of accepted proposals. Experiments conducted on both synthetic scenarios and protein design benchmarks validate the effectiveness of the approach. Ultimately, this research provides a statistically guaranteed adaptive termination mechanism for agent workflows, ensuring robust and reliable generation processes without compromising theoretical safety guarantees.
This study addresses the challenge of balancing false positive and false negative risks in AI fairness auditing under a specified tolerance threshold. To this end, it constructs a unified statistical testing framework tailored to two distinct objectives: violation certification and sensitivity screening. The proposed methodology introduces a constrained empirical likelihood test alongside an adaptive boundary surrogate principle. By integrating least favorable point calibration, split empirical likelihood, and spurious mark rate control strategies, the framework achieves differentiated trade-offs between error control and detection sensitivity. Numerical experiments validate the method’s capacity for flexible risk balancing, while an empirical analysis on the COMPAS dataset demonstrates its practical effectiveness in predictive fairness auditing.
This work addresses the challenge of efficiently verifying the validity of statistical learning models on a given data distribution without trusting the learner. It introduces publicly verifiable certificates of statistical validity (pvCSVs), establishing the first non-interactive, publicly verifiable proof system for learning that applies to adaptive statistical query (SQ) algorithms. The proposed framework enables any user to verify a model’s performance on their own distribution with a sample complexity of only $O(\log k)$, a significant improvement over the $\tilde{O}(\sqrt{k})$ complexity of standard SQ learning algorithms. This result provides a systematic characterization of the capabilities and limitations of the SQ model in the context of verifiable learning.