Score
Design and implement anytime‑valid sequential monitoring procedures (e‑processes) and tests—including betting‑based and sequential e‑testing constructions—that sequentially accumulate evidence and are calibrated on moderate‑deviation scales. Build conformance and drift monitors that control Type I error uniformly over time and support early stopping or decision rules when the accumulated e‑process crosses prescribed thresholds.
This work proposes a flexible and rigorous monitoring framework for two-arm randomized controlled trials that addresses the challenge of Type I error control in adaptive designs with frequent interim analyses and data-dependent adaptations. Built upon E-values and E-processes, the approach enables valid inference under composite null hypotheses and supports futility monitoring, seamlessly integrating group sequential and Bayesian perspectives. By constructing E-processes via betting martingales and incorporating calibration strategies, multiplicity adjustments, and hybrid design elements, the method guarantees strict Type I error control without requiring pre-specified analysis times. The framework is implemented in the open-source R package evalinger. Numerical experiments demonstrate that, under continuous monitoring, the proposed method not only maintains exact Type I error control but also achieves higher statistical power compared to conventional group sequential approaches.
In sequential anytime testing, combining e-process evidence across filtrations poses a fundamental challenge: an e-process valid under a coarse filtration may fail to retain validity under a finer one, even for the same null hypothesis. Method: We propose the “adjusted combination” framework, introducing the class of *adjuster* functions—characterizing necessary and sufficient conditions for cross-filtration evidence boosting. We prove that an adjuster is essential to restore anytime validity under the original filtration and quantify its logarithmic cost. Our approach unifies e-process theory, generalized test martingales, and filtration coarsening/refinement techniques. Results: We establish a complete characterization theorem for adjusters and validate the framework on real financial data for randomness testing. The work provides novel theoretical tools and practical methodology for sequential independence testing and predictive model evaluation.
This work addresses the problem of real-time, data-adaptive lower bounding of the number of true discoveries in online multiple testing, under the constraint that decisions must be based solely on past hypotheses and data—without lookahead. To tackle this challenge, we first establish that admissibility of online closed testing is equivalent to the availability of valid e-values at each time step. Building on this equivalence, we propose a novel online closed testing framework grounded in products of e-values, accompanied by an efficient algorithm and new e-value hedging and boosting mechanisms to enhance statistical power. Our method unifies modeling of exchangeable and arbitrarily dependent test statistics, ensuring strict control of both false discovery rate (FDR) and false discovery exceedance (FDX) under arbitrary dependence, while delivering tight lower confidence bounds on the number of true discoveries. Compared to state-of-the-art approaches, our method achieves theoretical unification and improved performance, enabling the first online procedure for true discovery control that simultaneously delivers strong empirical power and rigorous finite-sample guarantees.
Traditional fixed-sample hypothesis tests lack temporal flexibility—they cannot be terminated early or extended adaptively without inflating Type I error. Method: We propose an anytime-valid sequential testing framework that preserves statistical power. We rigorously prove that any fixed-sample test can be equivalently transformed into a sequential counterpart with identical power. This is achieved via p-value reconstruction and reinterpretation of significance levels, ensuring that the test can be stopped or continued at any time while strictly controlling the overall Type I error rate. Contributions/Results: We derive explicit anytime-valid versions of the z-test and t-test, which coincide exactly with their classical fixed-sample counterparts after N observations. We further show that the log-optimal sequential z-test corresponds to rejecting the null at the minimal future significance level required by the standard z-test. Our framework unifies fixed-sample and sequential paradigms, enabling reliable inference under dynamic, real-time data collection.
Existing e-BH procedures lack order-invariance over e-processes, causing test conclusions to reverse spuriously upon addition of irrelevant data and failing to control the false discovery rate (FDR) — or even the family-wise error rate (FWER) — under arbitrary dependence. This paper provides the first rigorous proof that e-BH violates FDR control in this setting. Method: We propose a novel, order-invariant multiple testing framework built on e-process upper bounds, featuring a dependence-structure-adaptive calibrator. Contribution/Results: Our method guarantees strict FDR control at level α (i.e., FDR-sup ≤ α) for arbitrary dependence structures among hypotheses. It eliminates temporal instability in rejection sets induced by sequential data arrival, ensuring robustness and reproducibility in dynamic data environments. Theoretical guarantees are established without restrictive assumptions on dependence, and the procedure is computationally tractable.
This study addresses the challenge of multiple testing in online change-point detection, where repeated hypothesis tests render traditional family-wise error rate (FWER) control inadequate. The work introduces the first formal definition of sequential FWER (sFWER) tailored to real-time data streams and proposes a simulation-based dynamic threshold calibration mechanism that effectively controls the probability of at least one false alarm within a sliding monitoring window. By integrating sequential inference with a moving window framework, the method operates without requiring a pre-specified number of tests and is well-suited for highly dependent data. Empirical evaluations demonstrate that the proposed approach accurately controls sFWER and substantially outperforms existing methods, with successful application to passive smartphone sensing data from adolescents exhibiting emotional instability.
This work addresses the challenge of rigorously validating whether a modified predictive distribution—arising, for instance, from calibration or intervention—is superior to the original predictor under arbitrary stopping times. The authors propose a general framework in which corrected distributions are constructed via non-negative predictable tilting, and a conditional e-value is derived from the likelihood ratio between the corrected and original predictions. This e-value is accumulated over time to form an e-process, enabling valid anytime-valid inference. The framework encompasses diverse correction types, including label shift and adjustments to conditional mean or variance, and incorporates conditional drift decomposition, rejection boundaries, and mixture strategies to ensure robust real-time validation. Both theoretical analysis and empirical experiments demonstrate that the method reliably detects improvements while strictly controlling Type-I error, making it suitable for dynamic and complex environments.
This study addresses critical limitations in current online A/B testing methodologies, where fixed-horizon tests suffer from inflated Type I error due to repeated monitoring, and prevailing sequential approaches struggle to simultaneously support futility stopping, control Type II error, and respect minimum detectable effect constraints. To overcome these challenges, the authors propose SPRT-z, a novel procedure grounded in Wald’s sequential probability ratio test. SPRT-z leverages large-sample normal approximations for computational efficiency, incorporates a scale-free horizon calibration (SFHC) mechanism to preserve statistical power under discrete monitoring, and employs a median-unbiased estimator derived from Brownian motion to correct bias induced by early stopping. Empirical evaluations demonstrate that the method rigorously controls both Type I and Type II errors, substantially reduces required sample sizes, and yields confidence intervals with coverage probabilities closely matching nominal levels.
研究解决了机器学习系统在线修正时监控器误报问题,通过Huber-style方法减少误报并提高有效性。
研究使用一致鞅方法进行分布无关的序列变化点检测,提出新的最优构造,并证明现有方法在控制PFA和ARL方面存在不足。