Equivalence testing with data-dependent and post-hoc equivalence margins

📅 2026-03-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional equivalence testing requires pre-specifying an equivalence margin, which is often difficult to determine objectively in practice. This work proposes a data-driven paradigm that leverages e-values to construct an equivalence margin guaranteed to cover the true effect with probability at least \(1 - \alpha\), and further generalizes this into a unified post-hoc selectable boundary curve. By abandoning fixed margins, the method applies to third-order strictly totally positive models—encompassing classical z- and t-tests—and yields boundaries with posteriorly valid coverage, offering stronger guarantees for decision-making. Compared to conventional fixed-margin approaches, the proposed framework enhances practical applicability and provides more informative guidance for inference.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationGame Theory and Economic Paradigms: EquilibriumReasoning under Uncertainty: Relational Probabilistic Models

Application Category

User Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating successEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
Equivalence testing compares the hypothesis that an effect $μ$ is large against the alternative that it is negligible. Here, `large' is classically expressed as being larger than some `equivalence margin' $Δ$. A longstanding problem is that this margin must be specified but can rarely be objectively justified in practice. We lay the foundation for an alternative paradigm, arguing to instead report a data-dependent margin $\widehatΔ_α$ that bounds the true effect $μ$ with probability $1 - α$. Our key argument is that $\widehatΔ_α$ is more useful than a test outcome at a fixed margin $Δ$, as measured by the guarantees it offers to decision makers. We generalize this to a curve of margins $α\mapsto \widehatΔ_α$, uniformly valid under the post-hoc selection of the margin. These ideas rely on e-values, which we derive for models that are strictly totally positive of order 3, nesting the classical z-test and t-test settings.
Problem

Research questions and friction points this paper is trying to address.

equivalence testing
equivalence margin
data-dependent
post-hoc
effect size
Innovation

Methods, ideas, or system contributions that make the work stand out.

equivalence testing
data-dependent margin
e-values
post-hoc inference
strict total positivity
🔎 Similar Papers
No similar papers found.
S
Stan Koobs
Econometric Institute, Erasmus University Rotterdam
N
Nick W. Koning
Econometric Institute, Erasmus University Rotterdam