probability of agreement

Designs and applies statistical models and inferential procedures to estimate or assess the probability that two measurements, raters, or procedures produce sufficiently similar values (i.e., are in agreement) within a prespecified tolerance; these procedures produce point estimates, confidence or credible intervals, and hypothesis tests for agreement probability while providing uncertainty quantification. This work includes building flexible measurement-error and dependence models and relaxing restrictive distributional assumptions so estimators and intervals remain valid under realistic error structures.

probabilityofagreement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.9
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a novel asymptotic framework for post-hoc valid inference that overcomes the rigidity of traditional statistical testing, which requires pre-specified significance levels. By extending e-value methodology from non-asymptotic to asymptotic settings, the paper enables the construction of confidence sets and p-values at any significance level chosen after observing the data. The approach achieves sharper and more efficient inference than existing non-asymptotic methods under weaker assumptions—specifically, without requiring strong moment conditions—thereby circumventing the conservativeness inherent in conventional procedures. This advancement establishes a rigorous theoretical foundation for flexible, data-driven statistical analysis while preserving inferential validity in large-sample regimes.

asymptotice-valueslarge-sample

This study addresses the critical need in clinical practice to assess whether new and existing measurement methods are interchangeable, which hinges on determining whether their results are clinically indistinguishable. To overcome the restrictive assumptions of current approaches—such as specific data distributions, homoscedasticity, and linear bias—the authors propose a more flexible inferential framework based on the Probability of Agreement (PoA). This framework integrates probabilistic modeling, statistical inference, and Monte Carlo simulation, thereby accommodating a broader range of real-world scenarios. The method is successfully demonstrated in a case study comparing tPSA measurement techniques and validated through extensive simulations, which confirm its robustness and superior performance. These advances substantially enhance the practical utility and generalizability of PoA-based interchangeability assessment.

clinical equivalenceflexible inferencemeasurement agreement

Tolerance Intervals Using Dirichlet Processes

Dec 01, 2025
SC
Seokjun Choi
🏛️ University of California, Santa Cruz | Genentech, Inc.

In nonclinical drug development, constructing tolerance intervals faces the challenge of simultaneously addressing small sample sizes, unknown underlying distributions, and statistical robustness. To address this, we propose a Bayesian nonparametric method based on the Dirichlet process (DP). This work is the first to directly leverage the analytical quantile process of the DP for tolerance interval construction—eliminating reliance on parametric distributional assumptions (e.g., normality) while substantially improving estimation efficiency under small samples. The method integrates Bayesian nonparametric modeling with computationally efficient algorithms, accommodating typical pharmaceutical data such as potency measurements. Simulation studies and real-data applications demonstrate that our approach maintains high robustness across skewed, heavy-tailed, and mixture distributions; achieves accuracy comparable to state-of-the-art parametric methods; and delivers superior performance on actual pharmaceutical datasets.

Develops Bayesian nonparametric tolerance intervals using Dirichlet processesOvercomes limitations of small samples and distribution misspecificationProvides robust tolerance interval estimation for pharmaceutical potency data

Asymptotic and compound e-values: multiple testing and empirical Bayes

Sep 29, 2024
NI
Nikolaos Ignatiadis
🏛️ University of Chicago | University of Waterloo | Carnegie Mellon University

This paper addresses the theoretical disconnect between false discovery rate (FDR) control and empirical Bayes inference in multiple hypothesis testing. Method: We formally define and systematically develop a rigorous theoretical framework for composite p-values and composite e-values—including asymptotic and fully composite settings—and propose a log-optimal mixture likelihood ratio method for constructing composite e-values. We establish a deep connection between composite e-values and empirical Bayes estimation, prove that any FDR-controlling procedure is equivalent to the e-BH procedure, and derive separable log-optimal composite e-values under point nulls; for heteroscedastic multi-t tests, we construct practical approximate e-values. Contribution: Our work unifies the p-value and e-value literatures, enables e-value derandomization, ensures cross-trial composability, and provides a complete, compositional characterization of FDR control via e-values.

Connecting compound e-values to FDR control proceduresConstructing asymptotic compound e-values for multiple t-testsDefining compound p-values and e-values for multiple testing

Quantifying Uncertainty: All We Need is the Bootstrap?

Mar 29, 2024
UZ
Urvsa Zrimvsek
🏛️ University of Ljubljana

This paper addresses the complexity and high pedagogical/practical barriers associated with conventional uncertainty quantification methods—such as standard errors, confidence intervals, and hypothesis tests—in statistical inference. To evaluate the potential of nonparametric bootstrap as a unified alternative, we conduct a large-scale simulation study rigorously comparing single bootstrap, double bootstrap, and classical methods across multiple dimensions: sample size, confidence level, data-generating mechanisms, and statistical functionals. Results demonstrate that the double bootstrap consistently achieves superior coverage accuracy, stability, and robustness—particularly under small-sample and non-normal conditions—outperforming both classical approaches and the single bootstrap. We thus establish the double bootstrap as a principled, parsimonious, and high-performance paradigm for uncertainty quantification, providing both theoretical justification and empirical evidence to support its adoption in statistical education and applied practice.

Assessing bootstrap's potential to simplify statistical education and practiceComparing double bootstrap performance against traditional confidence interval techniquesEvaluating bootstrap as universal alternative for uncertainty quantification methods

Latest Papers

What's happening recently
View more

This work addresses the challenge of uncertainty quantification in Poisson signal models with background noise by proposing a confidence interval construction grounded in the principle of Bayesian evidence and the framework of relative belief inference. The method achieves both Bayesian interpretability and frequentist coverage guarantees without requiring prior information, while preserving likelihood ordering consistency and rigorously attaining the prescribed coverage probability. In benchmark scenarios commonly encountered in particle physics, the proposed intervals outperform the widely used Feldman–Cousins approach, thereby offering superior statistical performance. Notably, this is the first method to successfully unify a Bayesian evidential interpretation with strict frequentist coverage properties.

confidencePoisson modelrelative belief

Traditional statistical inference relies on finite-dimensional modeling assumptions, which are often inadequate for uncertainty quantification in nonparametric function estimation. This work proposes a retrospective inference paradigm that centers on a point estimate derived from observed data and generates “estimation clones” to emulate repeated sampling. Inference is then conducted via the empirical distribution of these clones, without imposing probabilistic assumptions on the true parameter. The approach delivers robust uncertainty quantification in nonparametric regression while naturally accommodating parametric models, where it recovers classical inferential results. By integrating with smooth spline ANOVA models, the framework achieves both flexibility and theoretical rigor, offering a practical pathway for valid inference in nonparametric settings.

model assumptionsnonparametric estimationretrospective inference

This study addresses how to select an appropriate fusion rule for combining two probabilistic assessments of the same binary question, based on their semantic origins and dependency structure. It systematically analyzes the underlying assumptions of various probability fusion methods—such as averaging and odds multiplication—derives the corresponding combination formulas, and validates their performance under matched and mismatched conditions via Monte Carlo simulations. The work clarifies the prerequisite conditions for each fusion rule’s validity, demonstrates that relying solely on binary accuracy obscures differences in probability calibration, proposes a method to preserve shared premises for accurately computing success probabilities across multiple reasoning paths, and proves that pairwise fusion loses information when more than two paths are involved. Experiments confirm that correct application of fusion rules recovers true probabilities accurately, whereas misuse incurs substantial performance degradation.

conditional independenceevidence dependencepooling rules

This study addresses the estimation of the distribution function of a latent variable in an additive measurement error model, allowing the latent distribution to be arbitrary—discrete, continuous, or mixed—without requiring the existence of a density or global smoothness assumptions. Building upon Fourier inversion and the algebraic structure of the estimator introduced by Mynbaev et al., the authors develop a direct estimation approach that, for the first time, accommodates generalized distributions with multiple jump points. Theoretical analysis provides non-asymptotic bounds on bias and variance and establishes the asymptotic unbiasedness and consistency of the proposed estimator. Simulation studies demonstrate superior finite-sample performance compared to existing methods and confirm the practical feasibility of the recommended parameter selection scheme.

distribution functioninterval probabilityjump size

Hot Scholars

PW

Pascal Wullschleger

Dublin City University & Hochschule Luzern
Machine LearningDeep LearningNatural Language Processing
MP

Marc Pouly

Professor in Computer Science, Lucerne University of Applied Sciences and Arts
Machine LearningArtificial Intelligence
JF

Jennifer Foster

Hamilton Institute, Maynooth University
Natural Language ProcessingParsingSentiment AnalysisGrammar Checking
RD

Richard D Riley

University of Birmingham, UK.
Meta-analysisprognosis researchrisk prediction
GS

Gary S. Collins

Professor of Medical Statistics, University of Birmingham
medical statisticsstatisticsbiostatisticsmachine learning