Score
Applying statistical tests and procedures to determine whether observed effects are practically equivalent to a specified negligible magnitude, accounting for heterogeneity and uncertainty. Used to validate that verifications or detector differences are indistinguishable from no meaningful effect and to assess robustness of equivalence certificates.
Frequentist *p*-values and Bayesian posterior probabilities exhibit distinct operating characteristics in equivalence testing (e.g., bioequivalence), yet their comparative performance remains inadequately characterized. Method: We propose a novel Bayesian two-one-sided-tests (TOST) framework employing a uniform prior, the first to directly embed posterior probability into the TOST paradigm. We derive its theoretical relationship with frequentist *p*-values and quantify their alignment via evidence-measure correlation coefficients. The approach integrates power analysis, prior sensitivity assessment, and false discovery rate (FDR) control under multiple testing. Results: Simulation and theoretical analyses demonstrate that our method substantially improves statistical power—particularly under wide equivalence margins—while maintaining superior control of Type I error. It outperforms conventional *p*-value–based TOST in both single and multiple testing settings, offering enhanced statistical efficacy and more robust error regulation.
This paper addresses the inefficiency and subjectivity in sample size determination for Bayesian equivalence testing, where conventional methods rely on computationally intensive simulations and require ad hoc tuning of prior simulation counts or convergence thresholds. We propose a novel framework that controls sample size via the posterior Highest Density Interval (HDI) width, grounded in asymptotic normality theory for sample size. Specifically, we derive the first asymptotic normal approximation for HDI length and develop a two-stage numerical estimation procedure enabling efficient closed-form solutions under fixed-parameter models. Compared to standard Bayesian power-based approaches, our method accelerates computation by several-fold, eliminates user-specified simulation parameters, and guarantees that recommended sample sizes strictly satisfy prespecified statistical power and HDI precision constraints. The key contribution is the principled integration of HDI-width control with asymptotic theory—establishing a new paradigm for Bayesian sample size determination that achieves high accuracy, low computational cost, and complete parameter-freedom.
Existing methods for two-sample multi-quantile equivalence testing suffer from several limitations: mean-based approaches are inapplicable; tests for extreme quantiles are overly conservative; and statistical power is low under heteroscedasticity, especially with small or unbalanced samples. To address these challenges, this paper proposes the α-qTOST method—a quantile-based extension of the Two One-Sided Tests (TOST) framework—incorporating finite-sample correction to enable joint equivalence testing across multiple quantiles, including extreme ones. The method rigorously controls the nominal Type I error rate while substantially improving statistical power. Theoretical analysis and extensive simulations demonstrate its valid inference properties under heteroscedastic Gaussian models. The α-qTOST method has been successfully applied in HIV drug bridging trials and local drug delivery distribution assessments, providing robust support for equivalence decisions across diverse populations and operational conditions.
This paper addresses the problem of testing parameter homogeneity across multiple high-dimensional populations (with large $k$) under small per-group sample sizes. We propose an asymptotically distribution-free nonparametric test that relaxes the classical assumptions of fixed $k$ and large overall sample size. The test statistic is constructed via asymptotic distribution theory and enhanced by linear bootstrap calibration to improve finite-sample accuracy. We establish its asymptotic exactness under the null hypothesis and consistency under alternatives. Simulation studies demonstrate superior finite-sample performance compared to existing methods, and real-data analyses confirm its robustness and practical utility in high-dimensional, multi-group settings. The key contribution is the first nonparametric homogeneity test framework specifically designed for large-$k$, small-$n$ scenarios—offering both rigorous theoretical guarantees (asymptotic validity and consistency) and strong empirical performance.
Traditional chi-square tests can only assess exact independence between variables in contingency tables and are ill-suited for capturing the approximate independence commonly encountered in practice. This work proposes an equivalence testing framework tailored for two-dimensional contingency tables, which employs a boundary-point estimator combined with asymptotic critical values and Bootstrap resampling to enhance performance in small samples while maintaining statistical efficiency. The method is the first to be systematically applicable across contingency tables of varying dimensions, offering both theoretical rigor and computational feasibility. Extensive simulations demonstrate its superior performance across diverse table sizes, and its practical utility is further corroborated through real-data applications. The accompanying implementation code has been made publicly available.
Traditional significance testing often fosters dichotomous thinking, misinterpretation, and irreproducibility, thereby undermining robust scientific inference. This work proposes a unified evidence- and decision-oriented inferential framework that integrates advances in statistical inference and open science practices from 2016 to 2026. Moving beyond sole reliance on p-values, the framework incorporates compatibility interpretations, S-values, smallest effect sizes of interest (SESOI) equivalence testing, Bayesian workflows, and e-value sequential inference. It is further embedded within open science infrastructure—including preregistration, registered reports, multiverse analyses, and adherence to PRISMA 2020 and CONSORT 2025 reporting standards. By synergistically advancing methodological innovation and institutional reform, this open science inference system substantially enhances research transparency, reproducibility, and the quality of scientific decision-making.
This study addresses critical limitations in traditional pesticide risk assessment, which relies on a “no-effect” null hypothesis and suffers from insufficient statistical power—particularly in field trials with honey bees—leading to potential approval of high-risk compounds that fail to meet regulatory detection requirements. The authors introduce an equivalence testing framework that directly compares pesticide effects against predefined protection goals (e.g., ≤10% reduction in colony size) and evaluate its performance via Monte Carlo simulations. They validate the reliability of the EFSA-recommended approach in controlling false classifications of low risk and innovatively propose covariate adjustment and anti-clustering randomization strategies that significantly reduce the required number of trial sites while maintaining statistical power. Results underscore the necessity of increased replication for reliable assessment and demonstrate that, for pesticides with effects ≤5%, the proposed method reduces resource demands below current EFSA requirements, accompanied by an R toolkit and practical implementation guide.
This study addresses the limitations of current replication research, which relies on binary judgments to estimate replicability rates yet struggles to reliably distinguish true replicability due to inexact experimental replications and the absence of a shared data-generating mechanism. The authors propose two formal modeling frameworks—shared latent variables (as a benchmark) and conditional independence (as an operationalization)—to characterize statistical heterogeneity in inexact replications and its impact on replicability estimation. Through Bayesian hierarchical modeling, identifiability analysis, and quantification of heterogeneity, coupled with a reanalysis of the Many Labs 4 dataset, they demonstrate that standard methods fail to account for between-study heterogeneity, resulting in an irreducible lower bound on the variance of replicability estimates and a systematic underestimation of uncertainty. Consequently, aggregated replicability rates across heterogeneous studies lack stable interpretation, and conventional replication approaches are insufficient to reliably substantiate claims of a “replication crisis.”
This study addresses the problem of assessing whether observed data are “sufficiently close” to a binary generalized linear model—such as logistic regression—with fully categorical covariates, rather than requiring exact model fit. To this end, the authors propose a formal equivalence testing framework based on minimum distance methodology. The approach leverages both asymptotic theory and bootstrap procedures to compute critical values, thereby filling a critical gap left by conventional goodness-of-fit tests, which are ill-suited for evaluating practical equivalence. Through extensive simulation studies and analyses of two real-world datasets, the proposed method demonstrates strong finite-sample performance and practical utility, offering a robust tool for model adequacy assessment in applied settings.
This study addresses the challenge of disentangling sources of inter-laboratory variability—specifically baseline offsets versus differences in sensitivity—in multi-laboratory assessments of linear dose–response relationships. To this end, the authors propose a precision evaluation framework based on linear mixed-effects models, integrating analysis of variance, F-tests, and ISO 5725 standards to define and estimate repeatability and between-laboratory variance components. Overall measurement precision is quantified via average dose-specific variance. Under a fully balanced design, the framework yields an exact decomposition of total sum of squares and closed-form ANOVA estimators, overcoming the limitation of conventional fixed-effects models that detect only the presence of differences without identifying their origin. The approach was successfully applied to bronchoalveolar lavage fluid data from a rat intratracheal instillation study involving nanomaterials, effectively distinguishing the sources of observed variability.