Statistical Significance Revisited

📅 2026-05-07
📈 Citations: 0
Influential: 0
📄 PDF

career value

194K/year
🤖 AI Summary
This study addresses ongoing controversies surrounding null hypothesis significance testing—particularly the mechanical reliance on fixed p-value thresholds, rigid treatment of null and alternative hypotheses, and calls for abandoning the method altogether. By systematically evaluating recent reform proposals and integrating frequentist and Bayesian perspectives, the work re-examines the theoretical foundations and practical logic of significance testing. It clarifies the strengths and limitations of various reform approaches, demonstrates that sampling distributions can be constructed without pre-specifying thresholds or alternative hypotheses, and synthesizes techniques from hypothesis testing, Bayesian decision theory, confidence intervals, and sampling distribution analysis to expose key misinterpretations in current practice. The study argues that significance testing retains value for scientific inference when applied with greater caution and flexibility, rather than being discarded outright.
📝 Abstract
Since its introduction by Fisher, the method of hypothesis testing that relies on computing error probabilities has witnessed several developments. Perhaps the most significant development was the seminal contributions of Neyman and Pearson who brought in the concept of the alternative hypothesis with its corresponding error of the second kind. Significance tests have played a major role in various scientific and technological developments, but not without controversies. Although originally cast as frequentist approaches, Bayesian ideas have been incorporated into significance tests, widening access to them. The quantities central to computations of error probabilities are the sampling distributions, which can be computed even without thresholds or alternative hypotheses. Even though Fisher used the significance threshold of 0.05 in his calculations, he cautioned against prescribing any specific threshold. Recently, there have been calls for reformation in practice with regard to the almost standard use of the significance threshold of 0.05, prepublication confirmatory studies, the dichotomous consideration of the null and alternative hypothesis and abandoning significance tests altogether in favour of other approaches such as confidence intervals and Bayesian decision theory. In this paper, we examine these calls for reform and unearth their strengths and short comings.
Problem

Research questions and friction points this paper is trying to address.

statistical significance
hypothesis testing
p-value threshold
Neyman-Pearson framework
Bayesian methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

statistical significance
hypothesis testing
p-value threshold
Bayesian decision theory
sampling distributions