Score
Designs and analyzes statistical testing procedures that accumulate and evaluate evidence as data arrive, allowing for early stopping while provably controlling error rates (e.g., type I error) at arbitrary stopping times. This includes construction and calibration of sequential and two-sample tests, double-sequential sampling schemes, GLRT-based and likelihood-ratio aggregations, and anytime-valid e-value/e-process methods for online monitoring or finite-pool audits.
In sequential anytime testing, combining e-process evidence across filtrations poses a fundamental challenge: an e-process valid under a coarse filtration may fail to retain validity under a finer one, even for the same null hypothesis. Method: We propose the “adjusted combination” framework, introducing the class of *adjuster* functions—characterizing necessary and sufficient conditions for cross-filtration evidence boosting. We prove that an adjuster is essential to restore anytime validity under the original filtration and quantify its logarithmic cost. Our approach unifies e-process theory, generalized test martingales, and filtration coarsening/refinement techniques. Results: We establish a complete characterization theorem for adjusters and validate the framework on real financial data for randomness testing. The work provides novel theoretical tools and practical methodology for sequential independence testing and predictive model evaluation.
This work addresses the problem of real-time, data-adaptive lower bounding of the number of true discoveries in online multiple testing, under the constraint that decisions must be based solely on past hypotheses and data—without lookahead. To tackle this challenge, we first establish that admissibility of online closed testing is equivalent to the availability of valid e-values at each time step. Building on this equivalence, we propose a novel online closed testing framework grounded in products of e-values, accompanied by an efficient algorithm and new e-value hedging and boosting mechanisms to enhance statistical power. Our method unifies modeling of exchangeable and arbitrarily dependent test statistics, ensuring strict control of both false discovery rate (FDR) and false discovery exceedance (FDX) under arbitrary dependence, while delivering tight lower confidence bounds on the number of true discoveries. Compared to state-of-the-art approaches, our method achieves theoretical unification and improved performance, enabling the first online procedure for true discovery control that simultaneously delivers strong empirical power and rigorous finite-sample guarantees.
This work proposes a flexible and rigorous monitoring framework for two-arm randomized controlled trials that addresses the challenge of Type I error control in adaptive designs with frequent interim analyses and data-dependent adaptations. Built upon E-values and E-processes, the approach enables valid inference under composite null hypotheses and supports futility monitoring, seamlessly integrating group sequential and Bayesian perspectives. By constructing E-processes via betting martingales and incorporating calibration strategies, multiplicity adjustments, and hybrid design elements, the method guarantees strict Type I error control without requiring pre-specified analysis times. The framework is implemented in the open-source R package evalinger. Numerical experiments demonstrate that, under continuous monitoring, the proposed method not only maintains exact Type I error control but also achieves higher statistical power compared to conventional group sequential approaches.
Existing e-BH procedures lack order-invariance over e-processes, causing test conclusions to reverse spuriously upon addition of irrelevant data and failing to control the false discovery rate (FDR) — or even the family-wise error rate (FWER) — under arbitrary dependence. This paper provides the first rigorous proof that e-BH violates FDR control in this setting. Method: We propose a novel, order-invariant multiple testing framework built on e-process upper bounds, featuring a dependence-structure-adaptive calibrator. Contribution/Results: Our method guarantees strict FDR control at level α (i.e., FDR-sup ≤ α) for arbitrary dependence structures among hypotheses. It eliminates temporal instability in rejection sets induced by sequential data arrival, ensuring robustness and reproducibility in dynamic data environments. Theoretical guarantees are established without restrictive assumptions on dependence, and the procedure is computationally tractable.
Wald’s sequential probability ratio test (SPRT) suffers from overshoot at stopping times, preventing approximate thresholds—such as ((1-eta)/alpha) and (eta/(1-alpha))—from strictly controlling Type I/II error rates (when (eta > 0)) or guaranteeing optimality (when (eta = 0)). This paper introduces “sequential boosting”, a novel method that eliminates overshoot by constructing a corrected likelihood ratio statistic. It achieves, for the first time: (1) exact (alpha)/(eta) error control with strictly smaller expected sample size than approximate SPRT when (eta > 0); (2) optimal power-one performance—matching the theoretical lower bound on expected sample size—when (eta = 0); and (3) natural generalizations to confidence sequences, sampling-without-replacement settings, and conformal martingale frameworks. Theoretical analysis proves precise error calibration, while simulations demonstrate substantial sample-size reduction. The method is plug-and-play and broadly applicable across sequential inference paradigms.
This work addresses the problem of multiple hypothesis testing for edge distributions across multiple data streams. It proposes a sequential testing procedure that, for the first time, systematically incorporates arbitrary forms of prior information about the configuration of true and false hypotheses—such as known values or lower bounds on the number of active streams under each hypothesis, or mutual exclusivity constraints—while rigorously controlling the familywise error rate. By integrating sequential analysis with a search strategy over minimal alternative hypothesis configurations, the method achieves asymptotic optimality in terms of expected sample size among all valid procedures, without compromising reliability. Theoretical analysis establishes its computational efficiency and asymptotic optimality, and numerical experiments further demonstrate its substantial advantages in both testing efficiency and accuracy.
This study addresses the problem of determining whether high-frequency monitoring data return to their pre-intervention baseline distribution following an intervention. The authors propose a sequential testing procedure that requires no assumptions about the underlying data distribution. The method constructs a discrepancy measure via universal inference and combines it with individualized empirical calibration to form a non-negative supermartingale, yielding an e-process that enables valid detection of the recovery time at any arbitrary stopping point without specifying a null model. Theoretical analysis provides finite-sample bounds on the calibration error, and both simulations and a clinical case study demonstrate the method’s superior performance in accurately identifying the time at which baseline conditions are restored.
本文研究了通过自适应数据收集解决多重测试问题,提出基于e值的后验抽样方法(e-PS),有效控制错误发现率并减少样本需求。
This study addresses the localization of intervals containing false hypotheses in sequential data while controlling the false discovery rate (FDR). To this end, it proposes a sequential reset procedure based on e-values and test supermartingales, which optimizes the testing process by discarding low-information data. Furthermore, the authors construct anytime-valid FDR bounds that eliminate dependence on the testing horizon, providing rigorous theoretical guarantees under general dependency conditions. The primary contribution lies in achieving effective FDR control across both independent and dependent settings. Simulation studies validate the efficacy of the proposed approach, which is further demonstrated through successful applications to large language model watermark detection and financial backtesting tasks.
This study addresses a central challenge in statistical inference for clinical trials: achieving high operational flexibility—such as sample size re-estimation and treatment selection—while rigorously controlling the Type I error rate. The authors systematically integrate confirmatory adaptive designs with e-value-based anytime-valid testing methods, establishing for the first time their formal equivalence through conditional error functions and combination tests, while clarifying their distinct emphases on flexibility. The work constructs a theoretical bridge between these two frameworks: the e-value paradigm enhances optional continuation and loss control, whereas adaptive design principles can refine e-value testing strategies. This synthesis lays both a theoretical foundation and a practical pathway for developing next-generation inferential methods that simultaneously ensure strict error control and substantial procedural adaptability.