Score
Designs and applies statistical models and inferential procedures to estimate or assess the probability that two measurements, raters, or procedures produce sufficiently similar values (i.e., are in agreement) within a prespecified tolerance; these procedures produce point estimates, confidence or credible intervals, and hypothesis tests for agreement probability while providing uncertainty quantification. This work includes building flexible measurement-error and dependence models and relaxing restrictive distributional assumptions so estimators and intervals remain valid under realistic error structures.
This study addresses the evaluation of agreement among multiple measurement methods for continuous variables by systematically reviewing and synthesizing mainstream and emerging statistical approaches developed over the past two decades. It encompasses Bland–Altman analysis, Lin’s concordance correlation coefficient, and their extensions to robust, multivariate, repeated-measures, and spatial settings. Notably, the paper introduces probabilistic frameworks and spatial generalizations tailored to modern applications such as image analysis and environmental statistics. By clarifying the historical development, intrinsic connections, and limitations of these methods, the work establishes a unified methodological perspective and delineates promising directions for future research, thereby offering both theoretical grounding and practical guidance for selecting appropriate agreement assessment techniques.
This work proposes a novel asymptotic framework for post-hoc valid inference that overcomes the rigidity of traditional statistical testing, which requires pre-specified significance levels. By extending e-value methodology from non-asymptotic to asymptotic settings, the paper enables the construction of confidence sets and p-values at any significance level chosen after observing the data. The approach achieves sharper and more efficient inference than existing non-asymptotic methods under weaker assumptions—specifically, without requiring strong moment conditions—thereby circumventing the conservativeness inherent in conventional procedures. This advancement establishes a rigorous theoretical foundation for flexible, data-driven statistical analysis while preserving inferential validity in large-sample regimes.
This study addresses the critical need in clinical practice to assess whether new and existing measurement methods are interchangeable, which hinges on determining whether their results are clinically indistinguishable. To overcome the restrictive assumptions of current approaches—such as specific data distributions, homoscedasticity, and linear bias—the authors propose a more flexible inferential framework based on the Probability of Agreement (PoA). This framework integrates probabilistic modeling, statistical inference, and Monte Carlo simulation, thereby accommodating a broader range of real-world scenarios. The method is successfully demonstrated in a case study comparing tPSA measurement techniques and validated through extensive simulations, which confirm its robustness and superior performance. These advances substantially enhance the practical utility and generalizability of PoA-based interchangeability assessment.
In nonclinical drug development, constructing tolerance intervals faces the challenge of simultaneously addressing small sample sizes, unknown underlying distributions, and statistical robustness. To address this, we propose a Bayesian nonparametric method based on the Dirichlet process (DP). This work is the first to directly leverage the analytical quantile process of the DP for tolerance interval construction—eliminating reliance on parametric distributional assumptions (e.g., normality) while substantially improving estimation efficiency under small samples. The method integrates Bayesian nonparametric modeling with computationally efficient algorithms, accommodating typical pharmaceutical data such as potency measurements. Simulation studies and real-data applications demonstrate that our approach maintains high robustness across skewed, heavy-tailed, and mixture distributions; achieves accuracy comparable to state-of-the-art parametric methods; and delivers superior performance on actual pharmaceutical datasets.
This paper addresses the theoretical disconnect between false discovery rate (FDR) control and empirical Bayes inference in multiple hypothesis testing. Method: We formally define and systematically develop a rigorous theoretical framework for composite p-values and composite e-values—including asymptotic and fully composite settings—and propose a log-optimal mixture likelihood ratio method for constructing composite e-values. We establish a deep connection between composite e-values and empirical Bayes estimation, prove that any FDR-controlling procedure is equivalent to the e-BH procedure, and derive separable log-optimal composite e-values under point nulls; for heteroscedastic multi-t tests, we construct practical approximate e-values. Contribution: Our work unifies the p-value and e-value literatures, enables e-value derandomization, ensures cross-trial composability, and provides a complete, compositional characterization of FDR control via e-values.
This paper addresses the complexity and high pedagogical/practical barriers associated with conventional uncertainty quantification methods—such as standard errors, confidence intervals, and hypothesis tests—in statistical inference. To evaluate the potential of nonparametric bootstrap as a unified alternative, we conduct a large-scale simulation study rigorously comparing single bootstrap, double bootstrap, and classical methods across multiple dimensions: sample size, confidence level, data-generating mechanisms, and statistical functionals. Results demonstrate that the double bootstrap consistently achieves superior coverage accuracy, stability, and robustness—particularly under small-sample and non-normal conditions—outperforming both classical approaches and the single bootstrap. We thus establish the double bootstrap as a principled, parsimonious, and high-performance paradigm for uncertainty quantification, providing both theoretical justification and empirical evidence to support its adoption in statistical education and applied practice.
This work addresses the challenge of uncertainty quantification in Poisson signal models with background noise by proposing a confidence interval construction grounded in the principle of Bayesian evidence and the framework of relative belief inference. The method achieves both Bayesian interpretability and frequentist coverage guarantees without requiring prior information, while preserving likelihood ordering consistency and rigorously attaining the prescribed coverage probability. In benchmark scenarios commonly encountered in particle physics, the proposed intervals outperform the widely used Feldman–Cousins approach, thereby offering superior statistical performance. Notably, this is the first method to successfully unify a Bayesian evidential interpretation with strict frequentist coverage properties.
Traditional statistical inference relies on finite-dimensional modeling assumptions, which are often inadequate for uncertainty quantification in nonparametric function estimation. This work proposes a retrospective inference paradigm that centers on a point estimate derived from observed data and generates “estimation clones” to emulate repeated sampling. Inference is then conducted via the empirical distribution of these clones, without imposing probabilistic assumptions on the true parameter. The approach delivers robust uncertainty quantification in nonparametric regression while naturally accommodating parametric models, where it recovers classical inferential results. By integrating with smooth spline ANOVA models, the framework achieves both flexibility and theoretical rigor, offering a practical pathway for valid inference in nonparametric settings.
This study addresses how to select an appropriate fusion rule for combining two probabilistic assessments of the same binary question, based on their semantic origins and dependency structure. It systematically analyzes the underlying assumptions of various probability fusion methods—such as averaging and odds multiplication—derives the corresponding combination formulas, and validates their performance under matched and mismatched conditions via Monte Carlo simulations. The work clarifies the prerequisite conditions for each fusion rule’s validity, demonstrates that relying solely on binary accuracy obscures differences in probability calibration, proposes a method to preserve shared premises for accurately computing success probabilities across multiple reasoning paths, and proves that pairwise fusion loses information when more than two paths are involved. Experiments confirm that correct application of fusion rules recovers true probabilities accurately, whereas misuse incurs substantial performance degradation.
This study addresses the estimation of the distribution function of a latent variable in an additive measurement error model, allowing the latent distribution to be arbitrary—discrete, continuous, or mixed—without requiring the existence of a density or global smoothness assumptions. Building upon Fourier inversion and the algebraic structure of the estimator introduced by Mynbaev et al., the authors develop a direct estimation approach that, for the first time, accommodates generalized distributions with multiple jump points. Theoretical analysis provides non-asymptotic bounds on bias and variance and establishes the asymptotic unbiasedness and consistency of the proposed estimator. Simulation studies demonstrate superior finite-sample performance compared to existing methods and confirm the practical feasibility of the recommended parameter selection scheme.