Score
Designs and analyzes statistical testing procedures and test statistics that produce valid p-values and e-values, including construction and calibration of e-values and p-values, derivation of finite-sample error bounds (often via concentration inequalities), and formulation of test statistics that avoid ad-hoc tuning such as selecting a single ridge parameter. Builds analytic p-value combination and aggregation methods — including Cauchy-combination approaches — that combine dependent p-values, produce a single calibrated analytic p-value, maintain correct size across grids, and support testing of average, quantile, or information-type constraints.
This paper addresses the challenge of aggregating dependent p-values in multiple testing, focusing on heavy-tailed combination tests—such as the Cauchy and harmonic mean p-value methods—under asymptotically vanishing significance levels. We systematically characterize their statistical power under two asymptotic dependence regimes: (i) asymptotic independence, where they degenerate to Bonferroni correction, and (ii) quasi-asymptotic dependence, where they retain high power. Leveraging regular variation theory for tail distributions, coupled with asymptotic inference and multivariate t- or Gaussian dependence modeling, we rigorously establish the asymptotic validity of these tests. Monte Carlo simulations demonstrate that, particularly under strong dependence, these methods substantially outperform Bonferroni; notably, significant power gains emerge even at moderate significance levels (e.g., α = 0.01). Our work thus establishes a new paradigm for aggregating dependent p-values—one grounded in rigorous theory and offering practical advantages in real-world multiple testing scenarios.
This paper addresses the theoretical disconnect between false discovery rate (FDR) control and empirical Bayes inference in multiple hypothesis testing. Method: We formally define and systematically develop a rigorous theoretical framework for composite p-values and composite e-values—including asymptotic and fully composite settings—and propose a log-optimal mixture likelihood ratio method for constructing composite e-values. We establish a deep connection between composite e-values and empirical Bayes estimation, prove that any FDR-controlling procedure is equivalent to the e-BH procedure, and derive separable log-optimal composite e-values under point nulls; for heteroscedastic multi-t tests, we construct practical approximate e-values. Contribution: Our work unifies the p-value and e-value literatures, enables e-value derandomization, ensures cross-trial composability, and provides a complete, compositional characterization of FDR control via e-values.
This paper addresses the underutilization of heterogeneity and structural information in multiple hypothesis testing by proposing a general e-value–based framework. Methodologically: (1) it introduces a data-dependent weighting scheme—including a leave-one-out heuristic—for flexible aggregation of e-values across subsets, test statistics, and structure-informed covariates; (2) it unifies and extends the Benjamini–Hochberg (BH) and Benjamini–Yekutieli (BY) procedures to accommodate mixed tests and joint group-level–global false discovery rate (FDR) control; (3) it develops a structure-adaptive e-BH procedure that relaxes the independence and homogeneity assumptions inherent in classical p-value–based methods. Theoretically, it guarantees strict finite-sample FDR control. Numerical experiments demonstrate substantial gains in statistical power over state-of-the-art baselines—particularly under heterogeneous, grouped, or covariate-structured settings.
This paper addresses the limitations of traditional p-values in hypothesis testing by systematically establishing a theoretical framework and methodological system for e-values. Methodologically, it integrates and extends foundational theories—including universal inference, logarithmic optimality, e-processes, and multiple testing—leveraging probabilistic inequalities, sequential analysis, and information-theoretic optimality to develop novel methods for constructing, combining, and calibrating e-values. Key contributions include: (i) the first unified pedagogical and research paradigm for e-values; (ii) several unpublished theoretical advances, notably the optimality of e-processes in adaptive multiple testing; and (iii) a comprehensive textbook体系 tailored for graduate-level statistics education. Collectively, these results advance e-values as a cumulative, composable, and robust alternative to p-values for quantifying statistical evidence.
In resource-constrained multiple testing, exhaustive evaluation of all hypotheses and computation of exact test statistics (e.g., via experiments or precise calculations) is infeasible. Method: This paper proposes a surrogate-driven active testing framework that leverages auxiliary information—such as expert judgment, ML predictions, or historical data—to construct surrogate test statistics. It dynamically decides whether to invoke costly exact tests; otherwise, it substitutes the surrogate values directly. Contribution/Results: The framework is the first to enable compatible p-value and e-value constructions under arbitrary dependence structures—without requiring independence between surrogates and true statistics—while provably controlling the false discovery rate (FDR). By unifying active learning, multiple testing theory, and e-value theory, it achieves both theoretical rigor and practical utility. Empirical evaluation on scCRISPR causal effect analysis demonstrates a 32% increase in discoveries and a 68% reduction in computational cost compared to exhaustive testing, under identical FDR constraints.
This study addresses the challenge of effectively combining e-values under data-dependent tuning while preserving statistical validity. The authors develop a class of optimized e-process–based combination methods tailored for both independent settings and a newly introduced notion of “simultaneous e-variables.” They establish, for the first time, that combined tests retain validity even when tuning parameters are selected in a data-driven manner. By incorporating elementary symmetric polynomials, the proposed approach enhances statistical power and establishes a flexible intermediate framework that bridges independent and sequential validity. Through rigorous theoretical analysis grounded in dependence structure modeling, the method achieves substantially improved detection capability while maintaining strict Type-I error control.
论文提出Evidence-Calibration-Stability框架,通过区分证据、校准和稳定性来解决模型不确定性下的假设检验问题。
This study addresses the deviation of empirical probability integral transforms (PIT) from the theoretical uniform distribution under finite samples, a phenomenon induced by the two-stage sampling structure that invalidates conventional one-sample uniformity tests. The work systematically demonstrates that this non-uniformity arises from dependence structures and variance distortions introduced either by a reference sample or a rolling window: the former converges to a two-sample Kolmogorov–Smirnov distribution, while the latter exhibits temporal autocorrelation. Building on probability integral transform theory, empirical quantile estimation, and two-sample KS asymptotics, this paper establishes—for the first time—that empirical PIT values cannot be treated as independent uniform random variables. Leveraging these insights, the authors develop a corrected statistical inference framework specifically tailored for backtesting forecast calibration.
This work addresses the limitations of traditional conformal prediction, which relies on p-values and struggles to flexibly integrate multi-source evidence, while existing p-to-e calibration methods often distort original prediction sets, leading to excessive conservativeness. The paper proposes a set-preserving P2E calibrator that converts conformal p-values into e-values without altering the original prediction sets. This approach achieves, for the first time, a lossless transformation between p-values and e-values within conformal prediction, overcoming the conservativeness inherent in conventional calibration. It thereby enables effective e-value fusion and supports randomization techniques. Under strict 1−α coverage guarantees, the method significantly enhances the efficiency and accuracy of prediction sets in cross-conformal prediction and conformal aggregation, extending the theoretical framework for distribution-free uncertainty quantification.
本文研究了通过自适应数据收集解决多重测试问题,提出基于e值的后验抽样方法(e-PS),有效控制错误发现率并减少样本需求。