🤖 AI Summary
This study investigates the worst-case finite-sample distributional behavior of the log-likelihood ratio statistic in binary logistic regression without imposing regularity conditions on the design matrix. By leveraging extreme value analysis and concentration inequalities, together with a joint worst-case characterization over both the design vectors and the parameter space, the work establishes the first non-asymptotic, assumption-free uniform upper bound that parallels Wilks’ theorem. The main contributions include tight bounds on the (1−δ)-quantile: specifically, d log(en/d) + log(1/δ) in high dimensions, log log log n + log(1/δ) when d = 2, and a bound independent of the sample size n when d = 1. Under Gaussian designs, the analysis recovers the classical Wilks scaling while uncovering anomalous scaling laws in low dimensions.
📝 Abstract
We characterize the finite sample behavior of the log-likelihood ratio statistic in binary logistic regression, uniformly over both the design and the target parameter. For $n\geq d\geq 3$, we determine, up to universal constants, its worst case $(1-δ)$ quantile over all fixed collections of design vectors and all target parameters: \[ d\log\left(\frac{e n}{d}\right)+\log\left(\frac{1}δ\right). \] This is a nonasymptotic analogue of the Wilks $χ^2_d$ phenomenon and requires no regularity assumptions on the design. The low dimensional cases exhibit unusual behavior. The worst case quantile in dimension $d=2$ is sharply of order \[ \log\log\log n+\log\left(\frac{1}δ\right). \] The worst case quantile in dimension $d=1$ is of order $\log(1/δ)$, with no dependence on $n$. Finally, i.i.d. Gaussian design vectors recover the classical Wilks scale. In the regime $n\gtrsim d+\log(1/δ)$, we prove the sharp bound \[ d+\log\left(\frac{1}δ\right). \] Unlike existing asymptotic results, our bounds are uniform over the target parameter, which may depend on $n$, $d$, and $δ$.