Score
Design and implement empirical-likelihood-based estimators and inference procedures that replace parametric likelihoods with sample-based (nonparametric) likelihood constructions to produce point estimators, confidence regions, and hypothesis tests. Derive the asymptotic distributions of these estimators and test statistics, and develop bias-corrections and data-driven adjustments to ensure valid finite-sample inference for distributional features.
This paper addresses statistical inference under a setting featuring scarce labeled data, abundant unlabeled data, and high-quality external predictions. We propose a prediction-augmented empirical likelihood framework. Methodologically, we unify supervised estimating equations with auxiliary moment conditions derived from external predictions within an empirical likelihood optimization—learning the optimal auxiliary function via cross-fitting and basis expansion to yield a calibrated empirical distribution. Theoretically, our estimator achieves the semiparametric efficiency bound and enables valid estimation and uncertainty quantification for arbitrary differentiable functionals. Simulation and empirical studies demonstrate that the proposed method substantially reduces mean squared error, shortens confidence interval length, and strictly maintains nominal coverage—achieving both efficiency and robustness.
This study addresses the challenge of achieving both estimation efficiency and robustness in the presence of potential parametric model misspecification. The authors propose a hybrid likelihood approach that combines parametric and empirical likelihoods through a carefully designed weighting scheme, yielding an estimation framework that balances efficiency with robustness. They introduce a novel formulation of the hybrid likelihood function and establish the asymptotic normality of the resulting estimator along with a Wilks-type theorem, ensuring reliable performance even under model misspecification. Theoretical analysis demonstrates that the associated likelihood ratio statistic converges to a standard chi-squared distribution, providing a solid foundation for statistical inference. Additionally, a data-driven strategy is developed for selecting the tuning parameter that governs the trade-off between efficiency and robustness.
Nonprobability survey samples lack a rigorous sampling framework, impeding valid statistical inference. To address this, we propose a Pseudo-Empirical Likelihood (PEL) method for point and interval estimation of binary response variables. PEL preserves asymptotic unbiasedness and consistency while constructing range-preserving, data-driven confidence intervals—overcoming key limitations of conventional weighting and empirical likelihood approaches, particularly regarding coverage accuracy and boundary validity. Through theoretical derivation and extensive simulation studies, PEL demonstrates substantial improvements in coverage probability, average interval length, and stability, with especially pronounced advantages under severe sample representativeness bias. This work establishes a new inferential paradigm for complex nonprobability survey data that balances theoretical rigor with practical feasibility.
This study addresses the well-known issue that conventional likelihood ratio tests exhibit substantial size distortions under small samples when testing composite null hypotheses involving equality and inequality constraints, unions of multiple regions, or nuisance parameters, particularly due to violations of the regularity condition requiring a boundary-free manifold. To overcome these limitations, the authors propose a finite-sample testing procedure that conducts pointwise simple hypothesis tests across the entire null parameter space and applies a significance-level inflation correction to achieve overall size control. This approach accommodates generalized composite null structures beyond the scope of traditional methods. Numerical experiments demonstrate that the proposed test maintains actual size extremely close to the nominal level, with negligible distortion, in both small- and large-sample settings.
This paper addresses inference for parameters defined by multivariate three-sample U-statistics—such as the volume under the ROC surface (VUS) in multi-class classification—where conventional asymptotic approximations suffer from poor coverage and high computational cost. We propose a jackknife pseudo-value-based empirical likelihood (JEL) framework, the first to extend empirical likelihood to multivariate three-sample U-statistics. Under mild regularity conditions, the proposed JEL ratio statistic follows a Wilks-type chi-square asymptotic distribution, obviating explicit variance estimation and computationally intensive resampling. Monte Carlo simulations demonstrate that our method achieves coverage probabilities closer to nominal levels than normal approximation or kernel-based approaches, while reducing computation time substantially. Empirical evaluation on real classification datasets confirms its robustness and practical utility. Key contributions include: (i) theoretical innovation—a novel inferential framework with rigorous asymptotic theory; (ii) methodological advantages—variance-free and resampling-free inference; and (iii) applied impact—simultaneous improvements in accuracy and efficiency for VUS inference.
This study addresses Bayesian inference for low-dimensional target parameters in semiparametric models, particularly under the presence of complex nuisance components that may compromise frequentist properties. To this end, we construct posterior distributions by integrating estimating function methods with nonparametric Bayesian techniques—such as Dirichlet processes and Bayesian bootstrap—under conditions weaker than the classical stochastic equicontinuity assumption. We establish asymptotic normality and consistency of the resulting posterior, rigorously identifying the key assumptions required to guarantee desirable frequentist behavior. The theoretical analysis systematically elucidates how relaxing these assumptions affects inferential performance. Extensive simulations corroborate the effectiveness of the proposed methodology, demonstrating its robustness and accuracy in practical settings.
This work addresses the challenge that in many scientific applications, likelihood functions are only accessible implicitly through simulators, rendering traditional empirical Bayes methods—which rely on explicit likelihoods—ineffective. To overcome this limitation, the authors propose Simulation-Based Empirical Bayes (SBEB), the first approach to extend empirical Bayes to implicit-likelihood settings. SBEB integrates nonparametric empirical Bayes, simulation-based inference (SBI), and amortized inference networks, iteratively optimizing to learn a prior from observed data and simulator-generated samples that approximates the true population prior. Experimental results demonstrate that SBEB significantly outperforms conventional SBI methods employing fixed priors across multiple scientific simulators and real-world datasets, yielding substantially improved parameter estimation accuracy.
Traditional statistical inference relies on finite-dimensional modeling assumptions, which are often inadequate for uncertainty quantification in nonparametric function estimation. This work proposes a retrospective inference paradigm that centers on a point estimate derived from observed data and generates “estimation clones” to emulate repeated sampling. Inference is then conducted via the empirical distribution of these clones, without imposing probabilistic assumptions on the true parameter. The approach delivers robust uncertainty quantification in nonparametric regression while naturally accommodating parametric models, where it recovers classical inferential results. By integrating with smooth spline ANOVA models, the framework achieves both flexibility and theoretical rigor, offering a practical pathway for valid inference in nonparametric settings.
This work addresses the lack of a unified and efficient inference framework for models where exact likelihoods are intractable. It proposes the first general modeling and inference framework based on saddlepoint approximation, which preserves access to the moment-generating function through high-level operations to automatically construct the cumulant-generating function, its saddlepoint, and associated gradients. By integrating automatic differentiation, the framework optimizes the saddlepoint likelihood efficiently. It supports flexible modeling with complex distributional compositions and introduces diagnostic metrics that assess approximation error without requiring the true likelihood. Empirical results demonstrate that the method enables accurate parameter estimation, standard error computation, and error diagnosis with high computational efficiency and remarkable flexibility across multiple case studies.