The Sample Complexity of Distributionally Robust PAC Learning under Cressie--Read Divergences

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the sample complexity of distributionally robust PAC learning under Cressie–Read divergence constraints for binary classification with 0–1 loss, covering both realizable and agnostic settings. By analyzing the performance of empirical risk minimization under distributional perturbations and leveraging VC dimension theory, distributionally robust optimization, and refined concentration inequalities, the study extends existing results for χ² divergence to the entire Cressie–Read family with any order \(k > 1\). The derived sample complexity upper bounds are tight up to constant factors in the realizable case and up to logarithmic factors in the agnostic case, closing the gap between prior upper and lower bounds. The analysis further reveals an error amplification effect induced by robustness and a phase transition in the dependence on the divergence order; notably, as the perturbation radius vanishes, the classical PAC learning rates are precisely recovered.
📝 Abstract
We study distributionally robust PAC learning for the $0$--$1$-loss, where adversarial perturbations of the data distribution are constrained by a Cressie--Read divergence of order $k>1$ and radius $ρ\geq 0$. For hypothesis classes with VC dimension $d$, we establish realizable and agnostic sample-complexity bounds tight up to constant and logarithmic factors, respectively; ordinary empirical risk minimization attains both rates up to logarithmic factors. For target accuracy $\varepsilon\in(0,1)$ and confidence $δ\in(0,1)$, their respective orders are \[ \max\!\left\{\frac{1}{\varepsilon}, \frac{ρ^{\frac 1{k-1}}}{\varepsilon^{k_\star}} \right\}\cdot(d+\log δ^{-1}) \qquad\text{and}\qquad \max\!\left\{\frac{1}{\varepsilon^2}, \frac{ρ^{\frac1{k-1}}}{\varepsilon^{k_\star\vee 2}} \right\}\cdot(d+\log δ^{-1}), \] where $k_\star={k}/{(k-1)}$. For every fixed $ρ>0$, robustness changes the realizable $\varepsilon$-dependence from $\varepsilon^{-1}$ to $\varepsilon^{-k_\star}$ as $\varepsilon\downarrow0$. In the agnostic case, for $1<k<2$, robustness changes the $\varepsilon$-dependence from $\varepsilon^{-2}$ to $\varepsilon^{-k_\star}$, whereas for $k\geq2$ the exponent remains the classical $2$, with nontrivial $ρ$-dependence. Building on the known scalar reduction of robust $0$--$1$ risk to ordinary classification error, our analysis reveals a scale-sensitive interaction between the statistical estimation of classification error and its amplification by robustness, sharply explaining the transition in the agnostic rate. We extend the previously studied $χ^2$-divergence case to every Cressie--Read order $k>1$, close its upper--lower gaps, and recover standard PAC learning rates as $ρ\to0$, unlike previous bounds that fail to interpolate correctly in this limit.
Problem

Research questions and friction points this paper is trying to address.

distributionally robust learning
PAC learning
sample complexity
Cressie–Read divergence
0–1 loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

distributionally robust learning
Cressie–Read divergence
sample complexity
PAC learning
empirical risk minimization