🤖 AI Summary
This study addresses the challenges of inaccurate parameter estimation and inadequate goodness-of-fit testing for the Frank copula in small-sample settings. To this end, the authors propose a Bayesian estimation approach based on Jeffreys’ prior and conduct a systematic comparison with maximum likelihood estimation. The results demonstrate that, for sample sizes of 25 or fewer, the Jeffreys-prior Bayesian estimator substantially outperforms existing methods in terms of mean squared error. Furthermore, the work uncovers counterintuitive behaviors of commonly used goodness-of-fit test statistics and provides reliable Monte Carlo–based critical value tables. The proposed methodology is successfully applied to model the dependence between groundwater arsenic concentrations and hydrochemical variables in Đồng Tháp Province, Vietnam, thereby enhancing the practical implementation of copula models in small-sample scenarios.
📝 Abstract
This work has two major parts. First, we extend the recent study of Pham et al. (2025) on point estimation of the association parameter of a bivariate Frank copula. We investigate two Bayes estimators under the generalized flat prior and the Jeffreys prior, and compare them with the maximum likelihood estimator (MLE). Simulation results show that, for small sample sizes (n <= 25), the Bayes estimator under the Jeffreys prior uniformly outperforms both the generalized flat prior estimator and the MLE in terms of mean squared error (MSE). For moderate and large sample sizes, all estimators have very similar performances in terms of bias and MSE. We also discuss computational issues in the R package implementation that may significantly affect the computation of the MLE for very small samples.
In the second part, we apply the Frank copula to analyze the association between groundwater arsenic concentration and other hydrochemical variables using a recent dataset from Vietnam. We revisit the goodness-of-fit tests proposed by Genest et al. (2006), investigate several non-intuitive behaviors of the test statistics, and provide extensive simulated critical value tables. Our results complement and refine the computational findings reported in the earlier literature.