Score
Designs and applies statistical procedures to decide whether differences between models, systems, or treatments are practically negligible by testing whether observed gaps lie within pre-specified equivalence margins. This competence includes constructing and running tests and confidence-interval methods (e.g., two one-sided tests/TOST, high-confidence bounds, polynomial equivalence tests) and empirically assessing observational equivalence and behavioral divergence to quantify and bound negligible differences.
Traditional equivalence testing requires pre-specifying an equivalence margin, which is often difficult to determine objectively in practice. This work proposes a data-driven paradigm that leverages e-values to construct an equivalence margin guaranteed to cover the true effect with probability at least \(1 - \alpha\), and further generalizes this into a unified post-hoc selectable boundary curve. By abandoning fixed margins, the method applies to third-order strictly totally positive models—encompassing classical z- and t-tests—and yields boundaries with posteriorly valid coverage, offering stronger guarantees for decision-making. Compared to conventional fixed-margin approaches, the proposed framework enhances practical applicability and provides more informative guidance for inference.
Traditional chi-square tests can only assess exact independence between variables in contingency tables and are ill-suited for capturing the approximate independence commonly encountered in practice. This work proposes an equivalence testing framework tailored for two-dimensional contingency tables, which employs a boundary-point estimator combined with asymptotic critical values and Bootstrap resampling to enhance performance in small samples while maintaining statistical efficiency. The method is the first to be systematically applicable across contingency tables of varying dimensions, offering both theoretical rigor and computational feasibility. Extensive simulations demonstrate its superior performance across diverse table sizes, and its practical utility is further corroborated through real-data applications. The accompanying implementation code has been made publicly available.
This paper addresses the longstanding challenge of verifying the parallel trends assumption in Difference-in-Differences (DID) estimation, responding systematically to scholarly critiques of pre-treatment placebo tests. We propose a “conditional extrapolation assumption” framework, elevating pre-treatment testing from a mere screening device to a necessary condition for causal identification: treatment-effect estimation is only justified when pre-treatment trends exhibit no statistically significant deviation. Building on this, we construct confidence intervals with guaranteed conditional coverage, resolving the well-known post-selection inference failure of conventional pre-test–based approaches. We establish theoretical consistency and asymptotic validity of the method and demonstrate its finite-sample robustness and high coverage accuracy through simulations and empirical applications. Our core contribution lies in unifying pre-treatment testing with conditional inference, thereby substantially enhancing the credibility and reproducibility of DID-based causal identification.
Standard parallel-trends tests in Difference-in-Differences (DID) estimation merely fail to reject the null of no pre-treatment differences, but cannot actively confirm the assumption’s validity, thereby limiting causal credibility. This paper proposes the first equivalence-testing framework for DID pre-trend assessment: the null hypothesis posits *substantive* pre-treatment trend divergence between treatment and control groups; rejection of this null provides direct statistical support for the absence of meaningful pre-trends. We construct an asymptotically t-distributed test statistic and associated confidence sets within a two-way fixed-effects setting, ensuring theoretical rigor and empirical feasibility. The method naturally accommodates staggered adoption designs. Empirically, it detects subtle yet systematic pre-trend deviations missed by conventional tests—enhancing robustness, reproducibility, and causal interpretability of estimated treatment effects.
This paper addresses the inefficiency and subjectivity in sample size determination for Bayesian equivalence testing, where conventional methods rely on computationally intensive simulations and require ad hoc tuning of prior simulation counts or convergence thresholds. We propose a novel framework that controls sample size via the posterior Highest Density Interval (HDI) width, grounded in asymptotic normality theory for sample size. Specifically, we derive the first asymptotic normal approximation for HDI length and develop a two-stage numerical estimation procedure enabling efficient closed-form solutions under fixed-parameter models. Compared to standard Bayesian power-based approaches, our method accelerates computation by several-fold, eliminates user-specified simulation parameters, and guarantees that recommended sample sizes strictly satisfy prespecified statistical power and HDI precision constraints. The key contribution is the principled integration of HDI-width control with asymptotic theory—establishing a new paradigm for Bayesian sample size determination that achieves high accuracy, low computational cost, and complete parameter-freedom.
This study addresses the problem of testing for the existence of non-negative solutions to a system of linear equations when all parameters, including slope coefficients, are unknown. The authors propose a novel sample-splitting test that, for the first time, characterizes the closure of the null hypothesis under total variation distance, eliminating the need for simulated critical values and enabling applicability in high-dimensional settings with rapidly growing numbers of variables. By integrating total variation distance, sample splitting, and asymptotic theory under weak identification conditions, the method demonstrates strong power in both theoretical analysis and simulations. It combines computational simplicity with high-dimensional scalability and can be employed to construct confidence sets for partially identified parameters in nonparametric instrumental variable models.
This study addresses the risk of individual privacy leakage in standard equivalence testing, particularly in high-sensitivity domains such as healthcare. It introduces, for the first time, a unified differentially private two one-sided tests (DP-TOST) framework that enables privacy-preserving inference for both means and proportions. To handle the intractable sampling distributions in small-sample settings, the method employs simulation-based calibration, while integrating careful privacy budget allocation with the TOST paradigm to rigorously control Type I error rates under strong differential privacy guarantees. Theoretical analysis and extensive experiments demonstrate that, under reasonable privacy budgets or sample sizes, the proposed approach achieves statistical power nearly matching that of non-private benchmarks. Its reliability and practical utility are further validated through both simulated data and real-world medical datasets.
This study addresses the challenge of leveraging historical control data in randomized controlled trials to mitigate insufficient current control samples, a practice often compromised by population distributional differences that introduce bias and by existing methods’ difficulty in simultaneously controlling Type I error and maintaining statistical power. To resolve this, the authors propose a novel post-test fusion framework: first, a kernel-based two-sample equivalence test using Maximum Mean Discrepancy (MMD) assesses whether historical and current controls can be safely pooled; if equivalence is established, inference proceeds via a hybrid approach combining elements of the bootstrap and permutation tests. This method rigorously controls Type I error while substantially enhancing statistical power, thereby enabling robust and efficient utilization of historical control data.
This study addresses the efficient computation of simultaneous confidence intervals under family-wise error rate (FWER) control in multiple hypothesis testing. By extending, for the first time, the concept of coherence from closed testing procedures to the partitioning principle, the authors establish a formal algorithmic equivalence between these two frameworks, thereby unifying the construction of multiple testing procedures and simultaneous confidence intervals. Leveraging this theoretical connection, they propose a computationally efficient and practically feasible algorithm for constructing simultaneous confidence intervals. The method’s validity and advantages are demonstrated through illustrative examples, highlighting its effectiveness in real-world applications.