Score
Design and execute statistical non-inferiority and equivalence tests that assess whether a new treatment, procedure, model, or metric is not worse than a reference by more than a prespecified margin; this work includes defining the non-inferiority margin, specifying hypotheses and test statistics, computing confidence intervals or p-values for differences (e.g., pass rates or cost differences), and drawing and reporting non-inferiority/equivalence conclusions.
Non-inferiority studies in non-randomized data are vulnerable to unmeasured confounding bias, and conventional E-values—designed for statistical null hypotheses—lack direct interpretability with respect to clinically meaningful non-inferiority margins. Method: We extend the E-value framework to prespecified clinical non-inferiority thresholds, leveraging Ding and VanderWeele’s bias factor model to define the “non-inferiority E-value”: the minimum strength of unmeasured confounding, on the risk ratio scale, required to shift the effect estimate across the clinical threshold. Our approach computes E-values separately for point estimates and confidence limits, enabling clinically interpretable sensitivity analysis. Results: Applied to four empirical studies, non-inferiority E-values ranged from 1.0 to 3.0, revealing marked differences in conclusion robustness across study designs. This provides a transparent, actionable tool for bias assessment in non-randomized non-inferiority inference.
Traditional equivalence testing requires pre-specifying an equivalence margin, which is often difficult to determine objectively in practice. This work proposes a data-driven paradigm that leverages e-values to construct an equivalence margin guaranteed to cover the true effect with probability at least \(1 - \alpha\), and further generalizes this into a unified post-hoc selectable boundary curve. By abandoning fixed margins, the method applies to third-order strictly totally positive models—encompassing classical z- and t-tests—and yields boundaries with posteriorly valid coverage, offering stronger guarantees for decision-making. Compared to conventional fixed-margin approaches, the proposed framework enhances practical applicability and provides more informative guidance for inference.
Current regulatory guidance for non-inferiority trials inadequately incorporates the ICH E9(R1) estimand framework, leading to non-inferiority margin selection that overlooks how strategies for handling intercurrent events affect treatment effect estimation. This study addresses this gap by simulating patient pathways in a weight management context to systematically evaluate how different intercurrent event handling strategies—and their frequencies—alter the estimand. By reconstructing historical trial data under varying estimands, the research assesses the dependence of the non-inferiority margin \(M_1\) on the specified estimand. It reveals, for the first time, a strict correspondence between the non-inferiority margin and the estimand, demonstrating that even when clinical questions appear similar, historical control treatment effects cannot be directly reused if estimands differ. The findings underscore that margin specification must align precisely with the target estimand.
This paper addresses the inefficiency and subjectivity in sample size determination for Bayesian equivalence testing, where conventional methods rely on computationally intensive simulations and require ad hoc tuning of prior simulation counts or convergence thresholds. We propose a novel framework that controls sample size via the posterior Highest Density Interval (HDI) width, grounded in asymptotic normality theory for sample size. Specifically, we derive the first asymptotic normal approximation for HDI length and develop a two-stage numerical estimation procedure enabling efficient closed-form solutions under fixed-parameter models. Compared to standard Bayesian power-based approaches, our method accelerates computation by several-fold, eliminates user-specified simulation parameters, and guarantees that recommended sample sizes strictly satisfy prespecified statistical power and HDI precision constraints. The key contribution is the principled integration of HDI-width control with asymptotic theory—establishing a new paradigm for Bayesian sample size determination that achieves high accuracy, low computational cost, and complete parameter-freedom.
In methodological comparative studies, algorithmic failures—such as non-convergence or absence of output—preclude performance evaluation, yet existing literature lacks standardized guidelines for handling such failures, often overlooking or misapplying failure mitigation strategies. Method: We systematically analyze failure causes and risks of improper handling, critically examine prevalent censoring and imputation strategies for their statistical biases, and propose the principle of “context-adapted failure fallback,” establishing a framework grounded in empirically feasible fallback mechanisms. Through statistical modeling, failure root-cause diagnosis, and cross-domain empirical analysis, we identify widespread deficiencies in published studies’ failure handling practices. Contribution/Results: Two representative case studies demonstrate that inappropriate failure handling significantly distorts method rankings and undermines conclusion validity. Our work bridges critical theoretical and practical gaps in the principled treatment of algorithmic failures in empirical methodology research.
Traditional chi-square tests can only assess exact independence between variables in contingency tables and are ill-suited for capturing the approximate independence commonly encountered in practice. This work proposes an equivalence testing framework tailored for two-dimensional contingency tables, which employs a boundary-point estimator combined with asymptotic critical values and Bootstrap resampling to enhance performance in small samples while maintaining statistical efficiency. The method is the first to be systematically applicable across contingency tables of varying dimensions, offering both theoretical rigor and computational feasibility. Extensive simulations demonstrate its superior performance across diverse table sizes, and its practical utility is further corroborated through real-data applications. The accompanying implementation code has been made publicly available.
Traditional goodness-of-fit tests struggle to distinguish between “no significant difference” and “practical equivalence,” as failure to reject the null hypothesis may merely reflect insufficient test power. This work proposes the first kernel-based framework for full-distribution equivalence testing, leveraging Kernel Stein Discrepancy (KSD) and Maximum Mean Discrepancy (MMD) to quantify the distance between distributions while incorporating a prespecified minimum equivalence margin. By employing asymptotic normal approximations and bootstrap procedures to compute critical values, the method overcomes the limitations of existing equivalence tests, which are typically confined to parametric models or specific moments. Numerical experiments demonstrate that the proposed approach reliably assesses whether two distributions are equivalent within the specified margin while effectively controlling both Type I and Type II error rates.
This study addresses the long-standing misuse of the Wilcoxon signed-rank test in information retrieval (IR) evaluation, which has led to inflated Type I error rates and unreliable statistical inferences. Through a systematic literature review, theoretical analysis, and empirical validation using TREC data, the authors demonstrate that this misapplication stems from a common misconception: the Wilcoxon test is not a universally safe nonparametric alternative to the t-test and is particularly ill-suited for typical IR experimental settings. The work clarifies that the test’s assumptions are violated by the inherent dependencies and discrete score distributions characteristic of IR evaluation. Consequently, the paper urges the IR community to abandon the Wilcoxon signed-rank test in favor of more appropriate statistical methods to enhance the rigor and reliability of experimental methodology.
This study addresses a key challenge in multi-hypothesis group sequential clinical trials: how to provide informative simultaneous confidence intervals for effective treatment effect estimation while rigorously controlling the family-wise error rate (FWER). The authors propose a novel group sequential testing strategy that bases decisions solely on repeated p-values from the current stage and dynamically enhances significance thresholds by integrating evidence accumulated in prior stages. For the first time, they extend informative simultaneous confidence interval methodology to a graphical group sequential framework, combining the Bonferroni closure principle with repeated p-value methods. An iterative algorithm is developed to compute testing boundaries, complemented by precision assessment criteria and a conservative median estimation technique. The resulting approach maintains strict FWER control while substantially improving statistical power, with only minimal power loss attributable to the confidence intervals, and supports dynamic updating at each interim analysis stage.