Score
Design and build confidence sequences—sequences of confidence intervals or sets (including anytime-valid confidence intervals, sequential confidence intervals, anytime-valid confidence sets, and time-uniform policy sets) that guarantee prescribed coverage uniformly over time and under arbitrary stopping rules. Analyze their statistical properties (coverage, time-uniformity, width), derive stopping times and gap-dependent sample-complexity bounds, and construct procedures that retain validity under sequential data collection scenarios such as unknown logging.
Traditional model confidence sets rely on the fixed-sample assumption, rendering them inadequate for continuous model selection and uncertainty quantification under dynamic data streams. To address this, we introduce the first sequential extension of model confidence sets, proposing a dynamic model screening framework grounded in e-processes, confidence sequences, and sequential hypothesis testing. Our method requires no prespecified stopping rule and delivers, at any time, a nonasymptotic, time-uniform confidence set that provably covers the true optimal model subset with guaranteed nominal coverage probability. Unlike static approaches, it substantially enhances statistical robustness and real-time adaptability in online model evaluation. The framework provides both theoretical guarantees and practical tools for trustworthy model selection in streaming data analysis.
This work addresses the conservatism arising from the infinite-time validity assumption in sequential inference by introducing a “confidence horizon” framework that constructs anytime-valid confidence sequences within a finite time boundary. By integrating group sequential methods with adaptive Neyman allocation, the framework enables early stopping under budgetary or ethical constraints while preserving inferential accuracy. Key contributions include the first incorporation of a finite-time horizon into the anytime-valid inference paradigm, the establishment of explicit connections to classical group sequential boundaries (e.g., Pocock and O’Brien–Fleming), and the derivation of closed-form asymptotic quantiles that circumvent repeated integration. This analytical advance substantially improves the computational efficiency of critical value calculation and enhances the precision of treatment effect estimation.
This study addresses the frequentist prohibition against assigning a probability to the coverage of a parameter by an observed confidence interval, which limits nuanced interpretation of coverage events. By embedding confidence interval construction within a unified probabilistic framework through thought experiments and formal modeling, the work introduces a coverage indicator variable and, from a multi-level conditional probability perspective, demonstrates the coherence of assigning intermediate probabilities to single-instance coverage events under specific regularity conditions. This approach transcends the strict behaviorist constraints traditionally imposed on confidence intervals, revealing a tension between the exclusive reliance on design-stage coverage probability and the definition of long-run error rates. The result is a more flexible and internally consistent theoretical foundation for interpreting confidence intervals.
This work addresses the interpretational difficulty in classical frequentist inference, where assigning a post-hoc coverage probability to a specific confidence interval is traditionally prohibited. The authors model the coverage event of a confidence interval as a Bernoulli random variable and treat the nominal confidence level \(1-\alpha\) as a probabilistic prediction for this event, evaluated via strictly proper scoring rules. Within a frequentist framework, they combine decision theory, pivotal quantities, and parameter-free statistics to prove that \(1-\alpha\) is the unique optimal constant prediction. Furthermore, in unbounded translation-invariant models, they construct improved predictive forms—conditioned on ancillary statistics such as relative interval width—that yield non-constant but superior conditional coverage probabilities. This approach preserves frequentist validity while offering a coherent post-hoc probabilistic interpretation of confidence intervals, thereby resolving the classic interpretational paradox.
Traditional fixed-sample hypothesis tests lack temporal flexibility—they cannot be terminated early or extended adaptively without inflating Type I error. Method: We propose an anytime-valid sequential testing framework that preserves statistical power. We rigorously prove that any fixed-sample test can be equivalently transformed into a sequential counterpart with identical power. This is achieved via p-value reconstruction and reinterpretation of significance levels, ensuring that the test can be stopped or continued at any time while strictly controlling the overall Type I error rate. Contributions/Results: We derive explicit anytime-valid versions of the z-test and t-test, which coincide exactly with their classical fixed-sample counterparts after N observations. We further show that the log-optimal sequential z-test corresponds to rejecting the null at the minimal future significance level required by the standard z-test. Our framework unifies fixed-sample and sequential paradigms, enabling reliable inference under dynamic, real-time data collection.
This study addresses the problem of uncontrolled confidence interval error rates in dynamic evaluation of model leaderboards caused by repeated testing or early stopping. We propose an anytime-valid confidence sequence method for full-model ranking that combines betting e-processes for pairwise comparisons with closed testing to integrate all possible orderings. This approach constructs, for the first time under score-dependent conditions, confidence sequences that permit arbitrary stopping rules without inflating error rates. Simulations and experiments on public datasets demonstrate that the proposed method supports real-time monitoring and adaptive stopping, significantly reducing computational costs while sacrificing only minimal statistical power compared to fixed-sample approaches.
为解决重复获取数据时经典置信区间产生矛盾推断的问题,提出了一种计算广义线性模型中回归系数混合置信序列的高效方法。
This study addresses the challenge of efficiently comparing and selecting among multiple adaptive prediction pipelines under hard coverage constraints. To this end, it proposes the CC-SMCS framework, which achieves sequential model selection by decoupling feasibility from optimality. Technically, the approach employs stochastic constrained argmin modeling, simultaneous martingale confidence sequences, and rectangular region projections, yielding closed-form rules with finite-sample guarantees without requiring stationarity assumptions. The proposed method contains all constrained optimal pipelines with high probability while supporting data-dependent stopping and delayed feedback scenarios. Furthermore, this work establishes an impossibility result for margin-safe certification, thereby providing a theoretical foundation for constrained online learning.
This work addresses the challenge of verifying Markov decision processes (MDPs) with unknown transition probabilities that exhibit both nondeterminism and probabilistic uncertainty. To tackle this problem, the authors propose an online statistical model checking method grounded in confidence sequences. By integrating dynamic sampling, statistical hypothesis testing, and a novel online confidence sequence construction, the approach effectively mitigates the conservativeness and inefficiency inherent in traditional union-bound techniques. The resulting specialized verification tool maintains rigorous reliability guarantees while achieving a dramatic reduction in sample complexity—requiring on average approximately 50 times fewer samples than the current state-of-the-art methods—thereby substantially enhancing both efficiency and practical applicability.
本文针对固定宽度顺序停止规则在无限方差或长程依赖情况下的失效问题,提出了一种基于联合泛函极限定理的方法,并引入了序列子抽样过程来解决。