🤖 AI Summary
This study addresses the localization of intervals containing false hypotheses in sequential data while controlling the false discovery rate (FDR). To this end, it proposes a sequential reset procedure based on e-values and test supermartingales, which optimizes the testing process by discarding low-information data. Furthermore, the authors construct anytime-valid FDR bounds that eliminate dependence on the testing horizon, providing rigorous theoretical guarantees under general dependency conditions. The primary contribution lies in achieving effective FDR control across both independent and dependent settings. Simulation studies validate the efficacy of the proposed approach, which is further demonstrated through successful applications to large language model watermark detection and financial backtesting tasks.
📝 Abstract
Data arrive sequentially, each associated with a null hypothesis. We develop testing procedures to locate intervals in which some null hypotheses fail with false discovery rate (FDR) control. The new procedures are called sequential resetting procedures, and they are based on e-values and test supermartingales. We also develop a refined version of the procedures by dropping less informative data points before the block minimum of the test supermartingale in each rejection block. These procedures have explicit FDR bounds under two settings: a classic setting of independence and the more general setting of possible dependence across null data and non-null data. These FDR bounds are independent of the testing horizon, and the general one has anytime validity, but it has an extra logarithm factor compared with the standard FDR level. We present simulation studies and data experiments with applications of sequential resetting procedures to LLM watermark detection and financial backtesting.