perform likelihood-ratio testing

Design, implement, and analyze hypothesis tests and decision rules that use likelihood-ratio statistics, including computing log-likelihood ratios, p-values, decision thresholds, and sequential detectors or membership tests based on LLRs. Extend these methods to estimate and interpret ratio measures (e.g., odds ratios), perform likelihood-ratio tests for specific parameters such as skewness while deriving valid inference from asymptotic or approximate distributions, and build or fit any auxiliary non-member or null models required for ratio estimation and black-box membership calling.

performlikelihood-ratiotesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the implementation challenges of the conditional likelihood ratio (CLR) test under weak identification, where minimizing a non-convex objective function complicates computation and discrete approximations may undermine theoretical validity. The authors systematically evaluate the impact of two approximation strategies—the conventional grid search and the polynomial-based global optimization method proposed by Moreira et al.—on the size and power of the CLR test. By implementing both approaches within a linear instrumental variables framework, they demonstrate for the first time that the grid method can induce substantial size distortions or power losses under certain designs, whereas the polynomial optimization approach reliably preserves the test’s theoretical consistency and maintains robust performance across diverse data-generating processes.

conditional likelihood ratio testnon-convex optimizationtest power

The distribution of calibrated likelihood functions on the probability-likelihood Aitchison simplex

Sep 03, 2025
PN
Paul-Gauthier Noé
🏛️ Centre National de la Recherche Scientifique (CNRS) | Aix-Marseille Université | Individual | Laboratoire d’Informatique d’Avignon (LIA) | Avignon Université

This work addresses the lack of likelihood calibration in multiple hypothesis testing. Methodologically, it extends the binary-log-likelihood-ratio (LLR), idempotency constraints, and distribution calibration framework to arbitrary finite-dimensional hypothesis spaces—marking the first such generalization. Leveraging Aitchison geometry, it introduces an isometric isometric log-ratio (ilr) transformation on the probability simplex to vectorize Bayesian updates; this naturally generalizes LLRs and weights-of-evidence into vector-valued forms and jointly integrates probability calibration with nonlinear discriminant analysis. The resulting framework rigorously preserves the statistical interpretability and calibration of likelihood functions, with learned discriminant components directly corresponding to calibrated multiclass likelihoods. Experiments demonstrate that the proposed model significantly improves classification reliability while enhancing decision transparency and uncertainty quantification.

Develops calibrated multiclass likelihood functions for interpretable machine learningExtends likelihood function calibration beyond binary hypothesesGeneralizes calibration and idempotence using Aitchison simplex geometry

Improving Wald's (approximate) sequential probability ratio test by avoiding overshoot

Oct 21, 2024
LF
Lasse Fischer
🏛️ University of Bremen | Carnegie Mellon University

Wald’s sequential probability ratio test (SPRT) suffers from overshoot at stopping times, preventing approximate thresholds—such as ((1-eta)/alpha) and (eta/(1-alpha))—from strictly controlling Type I/II error rates (when (eta > 0)) or guaranteeing optimality (when (eta = 0)). This paper introduces “sequential boosting”, a novel method that eliminates overshoot by constructing a corrected likelihood ratio statistic. It achieves, for the first time: (1) exact (alpha)/(eta) error control with strictly smaller expected sample size than approximate SPRT when (eta > 0); (2) optimal power-one performance—matching the theoretical lower bound on expected sample size—when (eta = 0); and (3) natural generalizations to confidence sequences, sampling-without-replacement settings, and conformal martingale frameworks. Theoretical analysis proves precise error calibration, while simulations demonstrate substantial sample-size reduction. The method is plug-and-play and broadly applicable across sequential inference paradigms.

Avoid overshoot in Wald's approximate SPRT to ensure error controlExtend method to confidence sequences and non-replacement samplingImprove power-one SPRT for simple or one-sided hypotheses

This paper addresses the challenge of analytically constructing the optimal ROC curve in binary hypothesis testing when prior distributions are unknown or intractable. We propose the first maximum likelihood estimator for the ROC curve based on observed likelihood ratio samples (MLE-ROC). Unlike conventional approaches, MLE-ROC operates directly on likelihood ratio samples without requiring knowledge of the underlying data distributions and exhibits strong convergence under the Lévy metric. We establish theoretical consistency of its AUC estimator and demonstrate—via simulations—that it significantly outperforms the empirical ROC estimator, especially in small-sample and highly imbalanced settings with sparse negative instances, reducing estimation error by over 40%. The key contribution is the first formal parameterization of the ROC curve in the likelihood ratio domain within a maximum likelihood estimation framework, accompanied by rigorous asymptotic statistical guarantees.

Comparing accuracy of maximum likelihood vs. empirical estimators.Deriving finite sample bounds for ROC curve estimators.Estimating optimal ROC curves from likelihood ratio samples.

This study addresses the theoretical gap between classical hypothesis testing with fixed significance levels and Bayesian methods, particularly in light of the Lindley paradox. By leveraging moderate deviation theory, the authors develop a unified Bayesian framework for hypothesis testing. Through Bayesian risk analysis and asymptotic expansions, they show that the optimal test threshold operates on the scale of √(log n / n), naturally yielding Jeffreys’ threshold, the BIC penalty term, and the Chernoff–Stein error exponent. This framework not only resolves the Lindley paradox but also extends Rubin’s (1965) program to modern settings such as high-dimensional sparse inference, goodness-of-fit testing, and model selection. Moreover, it establishes the superiority of Bayesian procedures over classical Neyman–Pearson tests in terms of statistical risk.

Bayes riskBayesian hypothesis testingLindley paradox

Latest Papers

What's happening recently
View more

This study addresses the evaluation and selection of source-level likelihood ratio (LR) systems for forensic evidence-to-reference comparison tasks by proposing an integrated analytical framework that balances performance and practical feasibility. The authors employ strictly proper scoring rules to quantify how effectively each system updates Bayesian prior odds and present the first systematic comparison among specific-source feature-based, common-source anchored, and unanchored score-based LR approaches. Their findings reveal that specific-source feature-based LRs achieve the highest performance but incur substantial experimental costs, whereas common-source feature-based methods offer strong discriminative power with significantly reduced implementation complexity. All LR systems substantially outperform a baseline relying solely on prior odds. This work thus provides both theoretical grounding and practical guidance for selecting LR systems in forensic practice.

likelihood-ratio systemsperformance vs. feasibilityscoring rules

This study addresses the well-known issue that conventional likelihood ratio tests exhibit substantial size distortions under small samples when testing composite null hypotheses involving equality and inequality constraints, unions of multiple regions, or nuisance parameters, particularly due to violations of the regularity condition requiring a boundary-free manifold. To overcome these limitations, the authors propose a finite-sample testing procedure that conducts pointwise simple hypothesis tests across the entire null parameter space and applies a significance-level inflation correction to achieve overall size control. This approach accommodates generalized composite null structures beyond the scope of traditional methods. Numerical experiments demonstrate that the proposed test maintains actual size extremely close to the nominal level, with negligible distortion, in both small- and large-sample settings.

composite null hypothesesfinite-sample testinglikelihood ratio test

This study addresses the inadequate estimation accuracy of the relative risk (RR), odds ratio (OR), and their logarithmic forms for rare binary attributes in two populations. To overcome this limitation, the authors propose a sequential equal-allocation sampling method that efficiently estimates these parameters while ensuring the relative mean squared error (for RR/OR) or mean squared error (for the log-transformed parameters) remains below a pre-specified threshold. Under rare or moderately rare event settings, the proposed estimator achieves performance approaching the Cramér–Rao lower bound, offering both high efficiency and rigorous error control. The key innovation lies in integrating a sequential sampling strategy with explicit mean squared error constraints, substantially enhancing the precision and reliability of RR and OR estimation in scenarios involving rare events.

mean-square errorodds ratiorare events

This study addresses the low power and poor stability of conventional Wald-type tests for interval-censored data in small-sample settings by proposing the first likelihood ratio test framework based on spline sieves. The method integrates spline sieve modeling with likelihood ratio testing and rigorously derives its asymptotic distribution, ensuring both theoretical validity and practical robustness. Simulation studies demonstrate that the proposed approach substantially improves statistical power while maintaining proper control of Type I error rates. Its applicability and advantages are further confirmed through analysis of real clinical data.

Cox proportional hazards modelinterval-censored datasmall samples

This paper addresses the ethical and statistical trade-off between patient safety and statistical efficiency in sequential hypothesis testing, specifically concerning excessive exposure to inferior treatments. We propose a likelihood-ratio-driven adaptive sequential testing framework that integrates the Sequential Probability Ratio Test (SPRT) with a dynamic treatment allocation mechanism. Crucially, we derive, for the first time, an explicit closed-form upper bound on the number of assignments to the inferior treatment, thereby rigorously guaranteeing its finite exposure. Theoretical analysis establishes asymptotic statistical efficiency—achieving the optimal sample size growth rate under both hypotheses. Numerical simulations demonstrate that, compared to classical SPRT, our method maintains high correct-decision probability while reducing inferior-treatment allocations by 30–60%. Our core contribution lies in unifying statistical power, patient safety, and ethical constraints into a single coherent framework—offering a theoretically rigorous yet practically implementable paradigm for sensitive applications such as clinical trials.

Analytically ensures finite exposure to less effective treatmentDerives closed-form expression for expected inferior treatment applicationsDynamically concentrates sampling on superior population while maintaining efficiency

Hot Scholars

AR

Aaditya Ramdas

Associate Professor (with tenure), Carnegie Mellon University
Machine LearningStatistics
SC

Seungjin Choi

Research Director of Machine Learning Lab, Intellicode
machine learning
JB

Jay Bartroff

Professor of Statistics and Data Sciences, University of Texas at Austin
StatisticsProbability
DB

Debabrota Basu

Faculty, Inria at University of Lille and CNRS (CRIStAL), ELLIS Scholar
Reinforcement LearningMulti-armed BanditsDifferential PrivacyFairness
JQ

Jing Qin

University of Southern Denmark
MathematicsStatistics