🤖 AI Summary
This study addresses critical limitations in traditional pesticide risk assessment, which relies on a “no-effect” null hypothesis and suffers from insufficient statistical power—particularly in field trials with honey bees—leading to potential approval of high-risk compounds that fail to meet regulatory detection requirements. The authors introduce an equivalence testing framework that directly compares pesticide effects against predefined protection goals (e.g., ≤10% reduction in colony size) and evaluate its performance via Monte Carlo simulations. They validate the reliability of the EFSA-recommended approach in controlling false classifications of low risk and innovatively propose covariate adjustment and anti-clustering randomization strategies that significantly reduce the required number of trial sites while maintaining statistical power. Results underscore the necessity of increased replication for reliable assessment and demonstrate that, for pesticides with effects ≤5%, the proposed method reduces resource demands below current EFSA requirements, accompanied by an R toolkit and practical implementation guide.
📝 Abstract
Harmful pesticide effects exceeding specific protection goals (SPG) may go undetected in underpowered experimental designs. Regulatory honeybee field studies have consistently failed to reach the statistical power required under European Food Safety Authority (EFSA) guidance, which may have caused approval of high-risk substances. Therefore, EFSA advised a shift from testing the null hypothesis of 'no effect' to equivalence testing. Under this approach, a pesticide is classified as 'low risk' if the null hypothesis that its effect exceeds the SPG can be rejected. For honeybees, the recommended SPG is a colony size reduction below 10%. Critics have argued that this framework requires excessive site replication to demonstrate pesticide safety and proposed an alternative equivalence test defining treatment effects relative to the lower bound of the 90%-control-group confidence interval. Using simulations mimicking a regulatory honeybee field study, we show that although the two equivalence tests share the same trade-off between false 'low-risk' and false 'high-risk' classifications, only EFSA's original recommendation reliably identifies pesticides with effects > SPG at alpha = 0.2. Our results show that increasing site replication beyond the current practice is unavoidable for a reliable regulatory assessment. However, for pesticides with effect sizes of 5% or less, site requirements remain lower than those implied by the power requirement of the former EFSA guidance. Moreover, covariate adjustment through a model term or balanced colony allocation using anticlustering randomisation can reduce site requirements without losing power and thus save costs. Finally, we provide guidance and R functions for anticlustering randomisation and equivalence testing for pesticide risk assessment.