statistical hypothesis testing

Designs, implements, and interprets statistical hypothesis tests and inferential procedures for comparing models, estimating effects, and assessing relationships; this includes formulating null and alternative hypotheses, selecting and computing appropriate test statistics and p-values, conducting parametric and nonparametric tests, performing regression-based tests and agreement analyses, and evaluating independence and correlation. It also includes planning and calculating power and sample sizes, assessing distributional assumptions (robustness, symmetry, multimodality), controlling multiple-testing error rates (e.g., FDR), reproducing test analyses, and reporting effect sizes and uncertainty to support generalization.

statisticalhypothesistesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.43
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

On the analysis of sequential designs without a specified number of observations

Jul 03, 2025
AK
Anna Klimova
🏛️ National Center for Tumor Diseases (NCT) | Technical University | Eötvös Loránd University

This paper addresses sequential experimentation with categorical response variables, where a key challenge arises from outcome-dependent treatment assignment: subsequent interventions are dynamically determined by prior outcomes, inducing dependent observations, missing Cartesian product structure, and random, unpredictable total sample size. To handle this nonstandard data structure, the authors introduce staged trees—a graphical model from algebraic statistics—providing a directed-tree-based parametrization that jointly encodes sequential decision rules and response dependencies. They derive closed-form properties of maximum likelihood estimators, construct valid test statistics, and establish their asymptotic distributions. Extensive simulations and analysis of real sequential clinical trial data demonstrate the method’s validity and robustness. This work extends the algebraic statistical paradigm for sequential experiments and establishes an interpretable, computationally tractable foundation for adaptive design with categorical outcomes.

Analyzing sequential experiments with categorical responsesInvestigating distributional assumptions for staged treesModeling data without Cartesian product structure

Strategic Hypothesis Testing

Aug 05, 2025
SH
Safwan Hossain
🏛️ Harvard University | MPI for Intelligent Systems

This paper studies strategic hypothesis testing within a principal–agent framework: the agent holds private beliefs about product efficacy and may manipulate submitted data to maximize expected payoff; the principal must design a p-value threshold to balance Type I and Type II error risks. Methodologically, it innovatively integrates game-theoretic reasoning with classical statistical hypothesis testing by imposing incentive-compatibility constraints. The analysis establishes that the optimal p-value threshold exhibits a monotonic, analytically tractable structure in the agent’s strategic behavior, yielding a closed-form solution. Theoretically, it demonstrates that regulators can endogenously mitigate strategic reporting by calibrating the critical p-value, thereby unifying statistical robustness with incentive compatibility. Empirical validation using FDA drug approval data confirms the model’s predictive power, providing regulators with an interpretable and computationally tractable framework for optimizing approval policies.

Balances false positives and negatives with p-value thresholdsExamines strategic agent behavior in hypothesis testingValidates model using drug approval data

This study addresses the challenge in multiple hypothesis testing where existing methods struggle with unknown and arbitrary dependence structures among p-values, thereby limiting predictive power analysis and sample size planning. The authors propose the first Bayesian predictive power framework that accommodates arbitrary dependence without requiring independence assumptions, while supporting control of either the family-wise error rate (FWER) or the false discovery rate (FDR). By incorporating prior distributions on effect sizes, a uniform prior on the correlation matrix, and p-value weighting, the approach effectively mitigates p-hacking bias. Inference is carried out via Bayesian simulation using an asymmetric multivariate normal mean-variance mixture distribution with a scale-matrix mixture and a Dirichlet process prior, implemented in the R package bnpMTP. Application to a reanalysis of p-values from a lead exposure study demonstrates more robust power estimation and bias assessment, offering a reliable foundation for future sample size determination.

arbitrary dependencemultiple testing procedurep-value weighting

Active multiple testing with proxy p-values and e-values

Feb 08, 2025
ZX
Ziyu Xu
🏛️ Carnegie Mellon University

In resource-constrained multiple testing, exhaustive evaluation of all hypotheses and computation of exact test statistics (e.g., via experiments or precise calculations) is infeasible. Method: This paper proposes a surrogate-driven active testing framework that leverages auxiliary information—such as expert judgment, ML predictions, or historical data—to construct surrogate test statistics. It dynamically decides whether to invoke costly exact tests; otherwise, it substitutes the surrogate values directly. Contribution/Results: The framework is the first to enable compatible p-value and e-value constructions under arbitrary dependence structures—without requiring independence between surrogates and true statistics—while provably controlling the false discovery rate (FDR). By unifying active learning, multiple testing theory, and e-value theory, it achieves both theoretical rigor and practical utility. Empirical evaluation on scCRISPR causal effect analysis demonstrates a 32% increase in discoveries and a 68% reduction in computational cost compared to exhaustive testing, under identical FDR constraints.

Develops active multiple testing using proxy p-values and e-valuesEnsures false discovery rate control while maintaining high powerSelects hypotheses to query based on proxy statistics to save resources

Assessing Inference Methods

Dec 18, 2019
BF
Bruno Ferman
🏛️ Sao Paulo School of Economics - FGV

This study addresses the uncontrolled false positive rates and misleading inferences arising from commonly used simulation methods in shift-share designs. We systematically evaluate prevailing inferential approaches in empirical research through a suite of multilevel simulation experiments. By comparing Monte Carlo analysis with counterfactual data-generating mechanisms, we uncover non-monotonic trade-offs among fidelity, sensitivity, and risk of misdirection across simulation designs. We propose a novel “progressive-fidelity simulation framework,” demonstrating that low-fidelity simulations suffice to expose fundamental inferential flaws, whereas high-fidelity simulations detect subtle, previously overlooked biases—substantially improving detection power. The framework balances interpretability and computational efficiency, offering a reproducible and scalable paradigm for assessing the robustness of causal inference methods.

Analyzing trade-offs in simulation-based inference assessmentsEvaluating reliability of inference methods for false-positive controlProposing alternatives to misleading shift-share design evaluations

Latest Papers

What's happening recently
View more

This study addresses a critical limitation in existing design-based simulations used to evaluate inference methods, which often overstate bias induced by spatial correlation due to unrealistic data-generating mechanisms. In particular, share-shift designs that fix outcomes and resample shocks conflate true treatment effects with error dependence structures, leading to misleading assessments. To remedy this, the paper proposes an improved simulation framework that more accurately models error dependence and avoids spurious entanglement between treatment effects and error terms, thereby better approximating real-world data-generating processes. Integrating resampling techniques with share-shift analysis, the proposed approach substantially enhances the reliability of inference evaluation across multiple empirical applications, underscoring the essential role of aligning simulation designs with genuine underlying mechanisms for valid inference assessment.

data-generating processdesign-based simulationsinference validity

This study addresses the challenge of accurately inferring the distribution of individual treatment effects—such as the proportion benefiting, the median effect, or the maximum impact—in randomized experiments, without suffering power loss due to suboptimal pre-specified test statistics. The authors propose an adaptive randomization test that combines multiple rank-based statistics, ensuring finite-sample validity without requiring prior knowledge of the optimal statistic. Innovatively integrating adaptive statistic combination with stratified weighting, the method effectively circumvents the power degradation typically induced by multiple comparison corrections and accommodates heterogeneous stratified experimental designs. In an empirical application to a teacher training program, the approach reveals that approximately half of the teachers experience significant benefits, demonstrating superior detection power and interpretability compared to conventional single rank-based tests.

distributional inferenceindividual treatment effectsrandomization tests

Multiple testing

Jun 25, 2026

This study addresses the inflation of false positive rates in multiple hypothesis testing by systematically reviewing error rate control criteria—such as the family-wise error rate (FWER) and the false discovery rate (FDR)—and integrating both classical and contemporary correction methods. The work provides reproducible implementations of these approaches in R, offering a theoretically rigorous yet practice-oriented resource for teaching and applied research. By unifying methodological foundations with hands-on computational examples, this contribution fills a critical gap in existing textbooks, which often lack comprehensive integration of techniques and practical guidance. The resulting framework serves as a complete and efficient reference for graduate-level instruction and real-world data analysis, enhancing both pedagogical clarity and analytical reliability in high-dimensional statistical inference.

error controlhypothesis testingmultiple testing

Traditional significance testing often fosters dichotomous thinking, misinterpretation, and irreproducibility, thereby undermining robust scientific inference. This work proposes a unified evidence- and decision-oriented inferential framework that integrates advances in statistical inference and open science practices from 2016 to 2026. Moving beyond sole reliance on p-values, the framework incorporates compatibility interpretations, S-values, smallest effect sizes of interest (SESOI) equivalence testing, Bayesian workflows, and e-value sequential inference. It is further embedded within open science infrastructure—including preregistration, registered reports, multiverse analyses, and adherence to PRISMA 2020 and CONSORT 2025 reporting standards. By synergistically advancing methodological innovation and institutional reform, this open science inference system substantially enhances research transparency, reproducibility, and the quality of scientific decision-making.

evidence evaluationnull hypothesis significance testingopen science

Hot Scholars

LJ

Lucas Janson

Associate Professor, Harvard University Department of Statistics
High-Dimensional InferenceStatistical Machine Learning
AR

Aaditya Ramdas

Associate Professor (with tenure), Carnegie Mellon University
Machine LearningStatistics
SP

Samuel Pawel

Epidemiology, Biostatistics and Prevention Institute, University of Zurich
StatisticsMeta-Research
ED

Edgar Dobriban

Statistics & Computer Science, University of Pennsylvania
StatisticsMachine LearningAI
XS

Xiaofeng Shao

Department of Statistics and Data Science, and Dept of Economics, Washington University in St Louis
EconometricsFunctional data analysisHigh-dimensional data analysisTime series analysis