🤖 AI Summary
The reproducibility crisis in psychology is partly attributable to uncorrected multiple comparisons, inflating false discovery rates and contributing to replication failures. This study provides the first systematic quantification of how multiple comparisons contributed to non-replication across 88 psychological studies and introduces TreeBH—a novel false discovery rate (FDR) control procedure designed for hierarchical hypothesis structures. Empirical evaluation shows that applying TreeBH renders 21 originally significant findings non-significant; of these, 20 were indeed not replicated in follow-up studies—accounting for 34% of all non-replicated results—while preserving 97% statistical power. This work establishes uncorrected multiple comparisons as a key driver of the reproducibility crisis and delivers the first theoretically rigorous, empirically feasible hierarchical multiple testing correction framework tailored to typical experimental designs in psychology.
📝 Abstract
The field of psychological sciences has been grappling with the replicability crisis. Various issues have been identified as potential sources of this problem. We bring to light a potential source that has largely been overlooked and demonstrate its significant contribution to the problem: the practice of multiple comparisons. We analyzed 88 papers from the Reproducibility Project in Psychology and found that multiple results are commonly reported in a single paper, ranging from 4 to 730 (M=77.7), without multiple comparison adjustments. We retroactively applied such an adjustment using a hierarchical FDR controlling procedure (TreeBH; Bogomolov et al., 2021). 21 of 88 results were deemed insignificant after adjustment. Twenty of these 21 results indeed failed to replicate, constituting over a third of the non-replicable findings, while maintaining 97% power. We propose that this should become a common practice as an essential means to increase replicability in experimental psychology.