conduct non-inferiority testing

Design and execute statistical non-inferiority and equivalence tests that assess whether a new treatment, procedure, model, or metric is not worse than a reference by more than a prespecified margin; this work includes defining the non-inferiority margin, specifying hypotheses and test statistics, computing confidence intervals or p-values for differences (e.g., pass rates or cost differences), and drawing and reporting non-inferiority/equivalence conclusions.

conductnon-inferioritytesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A sensitivity analysis for non-inferiority studies with non-randomised data

Nov 22, 2025
DK
Daijiro Kabata
🏛️ Kobe University | Clinical & Translational Research Center, Kobe University Hospital

Non-inferiority studies in non-randomized data are vulnerable to unmeasured confounding bias, and conventional E-values—designed for statistical null hypotheses—lack direct interpretability with respect to clinically meaningful non-inferiority margins. Method: We extend the E-value framework to prespecified clinical non-inferiority thresholds, leveraging Ding and VanderWeele’s bias factor model to define the “non-inferiority E-value”: the minimum strength of unmeasured confounding, on the risk ratio scale, required to shift the effect estimate across the clinical threshold. Our approach computes E-values separately for point estimates and confidence limits, enabling clinically interpretable sensitivity analysis. Results: Applied to four empirical studies, non-inferiority E-values ranged from 1.0 to 3.0, revealing marked differences in conclusion robustness across study designs. This provides a transparent, actionable tool for bias assessment in non-randomized non-inferiority inference.

Extends E-value framework to assess unmeasured confounding in non-inferiority studiesProvides transparent robustness assessment for non-randomized clinical trial designsQuantifies bias needed to shift confidence limits to clinical non-inferiority margins

Traditional equivalence testing requires pre-specifying an equivalence margin, which is often difficult to determine objectively in practice. This work proposes a data-driven paradigm that leverages e-values to construct an equivalence margin guaranteed to cover the true effect with probability at least \(1 - \alpha\), and further generalizes this into a unified post-hoc selectable boundary curve. By abandoning fixed margins, the method applies to third-order strictly totally positive models—encompassing classical z- and t-tests—and yields boundaries with posteriorly valid coverage, offering stronger guarantees for decision-making. Compared to conventional fixed-margin approaches, the proposed framework enhances practical applicability and provides more informative guidance for inference.

data-dependenteffect sizeequivalence margin

Current regulatory guidance for non-inferiority trials inadequately incorporates the ICH E9(R1) estimand framework, leading to non-inferiority margin selection that overlooks how strategies for handling intercurrent events affect treatment effect estimation. This study addresses this gap by simulating patient pathways in a weight management context to systematically evaluate how different intercurrent event handling strategies—and their frequencies—alter the estimand. By reconstructing historical trial data under varying estimands, the research assesses the dependence of the non-inferiority margin \(M_1\) on the specified estimand. It reveals, for the first time, a strict correspondence between the non-inferiority margin and the estimand, demonstrating that even when clinical questions appear similar, historical control treatment effects cannot be directly reused if estimands differ. The findings underscore that margin specification must align precisely with the target estimand.

assay sensitivityconstancy assumptionestimand

Fast Sample Size Determination for Bayesian Equivalence Tests

Jun 15, 2023
LH
Luke Hagar
🏛️ The University of Queensland | University of Waterloo

This paper addresses the inefficiency and subjectivity in sample size determination for Bayesian equivalence testing, where conventional methods rely on computationally intensive simulations and require ad hoc tuning of prior simulation counts or convergence thresholds. We propose a novel framework that controls sample size via the posterior Highest Density Interval (HDI) width, grounded in asymptotic normality theory for sample size. Specifically, we derive the first asymptotic normal approximation for HDI length and develop a two-stage numerical estimation procedure enabling efficient closed-form solutions under fixed-parameter models. Compared to standard Bayesian power-based approaches, our method accelerates computation by several-fold, eliminates user-specified simulation parameters, and guarantees that recommended sample sizes strictly satisfy prespecified statistical power and HDI precision constraints. The key contribution is the principled integration of HDI-width control with asymptotic theory—establishing a new paradigm for Bayesian sample size determination that achieves high accuracy, low computational cost, and complete parameter-freedom.

Addressing non-normality in moderate sample size scenariosDetermining minimal sample size for precise interval estimatesProviding unified precision criteria framework for Bayesian/frequentist designs

On the handling of method failure in comparison studies

Aug 21, 2024
MW
Milena Wunsch
🏛️ LMU Munich | Munich Center for Machine Learning | Department of Statistics | MRC Clinical Trials Unit | UCL

In methodological comparative studies, algorithmic failures—such as non-convergence or absence of output—preclude performance evaluation, yet existing literature lacks standardized guidelines for handling such failures, often overlooking or misapplying failure mitigation strategies. Method: We systematically analyze failure causes and risks of improper handling, critically examine prevalent censoring and imputation strategies for their statistical biases, and propose the principle of “context-adapted failure fallback,” establishing a framework grounded in empirically feasible fallback mechanisms. Through statistical modeling, failure root-cause diagnosis, and cross-domain empirical analysis, we identify widespread deficiencies in published studies’ failure handling practices. Contribution/Results: Two representative case studies demonstrate that inappropriate failure handling significantly distorts method rankings and undermines conclusion validity. Our work bridges critical theoretical and practical gaps in the principled treatment of algorithmic failures in empirical methodology research.

Addressing method failure handling in comparison studiesProviding guidance on proper failure interpretation and reportingRecommending realistic fallback strategies for method failures

Latest Papers

What's happening recently
View more

Traditional chi-square tests can only assess exact independence between variables in contingency tables and are ill-suited for capturing the approximate independence commonly encountered in practice. This work proposes an equivalence testing framework tailored for two-dimensional contingency tables, which employs a boundary-point estimator combined with asymptotic critical values and Bootstrap resampling to enhance performance in small samples while maintaining statistical efficiency. The method is the first to be systematically applicable across contingency tables of varying dimensions, offering both theoretical rigor and computational feasibility. Extensive simulations demonstrate its superior performance across diverse table sizes, and its practical utility is further corroborated through real-data applications. The accompanying implementation code has been made publicly available.

approximate independenceasymptoticbootstrap

Traditional goodness-of-fit tests struggle to distinguish between “no significant difference” and “practical equivalence,” as failure to reject the null hypothesis may merely reflect insufficient test power. This work proposes the first kernel-based framework for full-distribution equivalence testing, leveraging Kernel Stein Discrepancy (KSD) and Maximum Mean Discrepancy (MMD) to quantify the distance between distributions while incorporating a prespecified minimum equivalence margin. By employing asymptotic normal approximations and bootstrap procedures to compute critical values, the method overcomes the limitations of existing equivalence tests, which are typically confined to parametric models or specific moments. Numerical experiments demonstrate that the proposed approach reliably assesses whether two distributions are equivalent within the specified margin while effectively controlling both Type I and Type II error rates.

distributional equivalenceequivalence testingkernel methods

This study addresses the long-standing misuse of the Wilcoxon signed-rank test in information retrieval (IR) evaluation, which has led to inflated Type I error rates and unreliable statistical inferences. Through a systematic literature review, theoretical analysis, and empirical validation using TREC data, the authors demonstrate that this misapplication stems from a common misconception: the Wilcoxon test is not a universally safe nonparametric alternative to the t-test and is particularly ill-suited for typical IR experimental settings. The work clarifies that the test’s assumptions are violated by the inherent dependencies and discrete score distributions characteristic of IR evaluation. Consequently, the paper urges the IR community to abandon the Wilcoxon signed-rank test in favor of more appropriate statistical methods to enhance the rigor and reliability of experimental methodology.

Information Retrievalnon-parametric teststatistical misuse

This study addresses a key challenge in multi-hypothesis group sequential clinical trials: how to provide informative simultaneous confidence intervals for effective treatment effect estimation while rigorously controlling the family-wise error rate (FWER). The authors propose a novel group sequential testing strategy that bases decisions solely on repeated p-values from the current stage and dynamically enhances significance thresholds by integrating evidence accumulated in prior stages. For the first time, they extend informative simultaneous confidence interval methodology to a graphical group sequential framework, combining the Bonferroni closure principle with repeated p-value methods. An iterative algorithm is developed to compute testing boundaries, complemented by precision assessment criteria and a conservative median estimation technique. The resulting approach maintains strict FWER control while substantially improving statistical power, with only minimal power loss attributable to the confidence intervals, and supports dynamic updating at each interim analysis stage.

family-wise error rategroup sequential trialsmultiple hypotheses

Hot Scholars

MS

Mike Schaekermann

Computer Science PhD, Eng BSc, Medicine State Exam I
Human-Computer InteractionMachine LearningMedicine
WW

Wenjing Wu

Rice University
Two-dimensional materials