Score
Designs and performs analyses, models, and measurements that quantify and characterize the probability a statistical test or decision procedure will detect true effects (statistical power), including deriving local power, superiority/inferiority and equivalence conditions, and handling correlated or non‑Gaussian cases. Also develops and applies measurement methods, instrumentation, and resource trade‑off analyses (e.g., power–area or power-and-area tradeoffs) to evaluate and optimize detection performance and associated costs.
This paper addresses the problem of diminished statistical power in nonparametric tests arising from excessive dependence between test statistics and auxiliary statistics. To resolve this, we propose a novel framework grounded in statistical independence principles. Methodologically, we reformulate hypothesis testing via a relativity principle and establish asymptotic independence between test and auxiliary statistics—yielding a Basu-type theoretical guarantee—while preserving distributional invariance under both null and alternative hypotheses. Integrating decision-theoretic criteria with explicit independence constraints, our approach systematically enhances classical tests, including Shapiro–Wilk, Anderson–Darling, Kolmogorov–Smirnov, and symmetry-center tests. Extensive simulations demonstrate that the proposed methods significantly improve power over conventional approaches while retaining robustness and computational efficiency, making them broadly applicable across diverse nonparametric settings.
In online A/B testing, joint inference across multiple metrics faces a fundamental trade-off: metric aggregation incurs information loss, while stringent multiple-testing correction severely diminishes statistical power. This paper proposes the first Bayesian experimental design framework that jointly optimizes false discovery rate (FDR) control and average statistical power. Innovatively formulating FDR and power as dual constraints, our method efficiently computes the optimal sample size and decision thresholds via only two numerical simulations—bypassing the computational bottleneck of traditional intensive Monte Carlo simulation. Evaluated on real-world multi-metric scenarios, the approach achieves strict FDR control at ≤0.1 while improving average statistical power by 22% relative to standard methods. Moreover, it accelerates computation by over 90% compared to exhaustive grid search, enabling scalable, principled multi-metric experimentation.
Bayesian experimental design with nuisance parameters remains challenging due to the need to account for their prior uncertainty while maintaining statistical rigor and computational feasibility. Method: This paper proposes a fully Bayesian framework that explicitly models prior uncertainty over all parameters—including nuisance parameters—during the design stage. It introduces Bayesian additive regression trees (BART) to the experimental design literature for the first time, integrating asymptotic posterior approximations with Monte Carlo simulation to efficiently optimize sample size and decision rules under both fixed and adaptive designs. Contribution/Results: The approach significantly reduces computational burden compared to conventional resampling-intensive methods, while preserving statistical power and robust operating characteristics. Key innovations include: (1) unified quantification of nuisance parameter uncertainty; (2) a BART-driven design function learning mechanism that enhances interpretability and generalizability; and (3) an end-to-end Bayesian design pipeline balancing robustness, efficiency, and practical implementation.
This paper addresses the statistical analysis challenge of set-valued data (e.g., EMI injection point sets from electronic devices) in inter-laboratory comparisons. Methodologically, it proposes a consensus inference–oriented modeling framework that innovatively integrates Hamming distance to quantify set dissimilarity, Fisher’s noncentral hypergeometric distribution to model deviation counts, and a Bayesian hierarchical model to disentangle inter-laboratory consensus from intra-laboratory variability. Key contributions include: (i) the first application of the noncentral hypergeometric distribution to set-based consensus modeling, enabling statistically rigorous quantification of deviation counts; (ii) simultaneous estimation of a global consensus set and laboratory-specific offsets via hierarchical Bayesian inference; and (iii) substantially improved comparability and reliability of multi-laboratory results. The method is validated on real-world EMC inter-comparison data, demonstrating its effectiveness in identifying robust consensus sets and quantifying intra-laboratory variation.
In resource-constrained multiple testing, exhaustive evaluation of all hypotheses and computation of exact test statistics (e.g., via experiments or precise calculations) is infeasible. Method: This paper proposes a surrogate-driven active testing framework that leverages auxiliary information—such as expert judgment, ML predictions, or historical data—to construct surrogate test statistics. It dynamically decides whether to invoke costly exact tests; otherwise, it substitutes the surrogate values directly. Contribution/Results: The framework is the first to enable compatible p-value and e-value constructions under arbitrary dependence structures—without requiring independence between surrogates and true statistics—while provably controlling the false discovery rate (FDR). By unifying active learning, multiple testing theory, and e-value theory, it achieves both theoretical rigor and practical utility. Empirical evaluation on scCRISPR causal effect analysis demonstrates a 32% increase in discoveries and a 68% reduction in computational cost compared to exhaustive testing, under identical FDR constraints.
This work addresses the frequent lack of systematic and credible statistical evaluation in ECE/CS research, which often undermines the persuasiveness of empirical claims. To bridge this gap, we propose a structured statistical evaluation workflow tailored for beginners, integrating classical methods—such as t-tests and ANOVA—with modern nonparametric techniques, including bootstrap resampling, Wilcoxon tests, and Cliff’s delta. The framework spans the entire pipeline from formulating research claims to reporting results, supporting factorial designs, multiple comparison corrections, and simulation-based validation. Accompanying the methodology are fully reproducible Python implementations, illustrative examples, and a pre-submission checklist. This approach substantially enhances the reliability and reproducibility of experimental findings while offering both pedagogical utility and practical guidance for researchers.
In nonstandard testing scenarios, evaluating the optimality of heuristic test statistics is challenging. Method: This paper proposes a nested numerical optimization framework to assess whether a given heuristic test is approximately optimal—i.e., whether its power curve approximates the power envelope generated by a weighted average power (WAP)-maximizing test. Contribution/Results: We introduce, for the first time, a data-driven approach to approximate the optimal weighting function. Crucially, we establish theoretically that the rejection probability of the WAP-optimal test itself constitutes a tight upper bound on the power of any heuristic test—a result both theoretically unexpected and practically valuable. The method is provably convergent and is successfully applied to two canonical problems: robust conditional likelihood ratio (CLR) testing under weak instruments and testing at nuisance parameter boundaries. Empirical results confirm the near-optimality of the heuristic tests in these settings.
Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.
This study addresses the challenge in multiple hypothesis testing where existing methods struggle with unknown and arbitrary dependence structures among p-values, thereby limiting predictive power analysis and sample size planning. The authors propose the first Bayesian predictive power framework that accommodates arbitrary dependence without requiring independence assumptions, while supporting control of either the family-wise error rate (FWER) or the false discovery rate (FDR). By incorporating prior distributions on effect sizes, a uniform prior on the correlation matrix, and p-value weighting, the approach effectively mitigates p-hacking bias. Inference is carried out via Bayesian simulation using an asymmetric multivariate normal mean-variance mixture distribution with a scale-matrix mixture and a Dirichlet process prior, implemented in the R package bnpMTP. Application to a reanalysis of p-values from a lead exposure study demonstrates more robust power estimation and bias assessment, offering a reliable foundation for future sample size determination.