🤖 AI Summary
This study investigates the trade-off between robustness and generalization in adversarial training and the resulting statistical accuracy loss. Working within the reproducing kernel Hilbert space (RKHS) framework, the authors analyze the generalization error of adversarially trained estimators via integral operators and establish the first matching upper and lower generalization bounds, elucidating how the interplay between noise and robustness governs approximation rates. To mitigate this degradation, they propose a two-stage denoising-and-correction procedure that, under an appropriate level of robustness, recovers the minimax optimal polynomial convergence rate up to logarithmic factors. Both theoretical analysis and numerical experiments confirm the effectiveness of the proposed method.
📝 Abstract
Adversarial training has emerged as a powerful approach for protecting models against adversarial attacks in a broad range of real-world applications. In this paper, we study adversarial training in the reproducing kernel Hilbert space (RKHS) framework through the associated kernel integral operator. We first derive source-uniform generalization error bounds for the RKHS adversarial training estimator in terms of the robustness level, sample size, source smoothness, and kernel spectrum. On a fixed polynomial-spectrum model, we further establish a matching lower bound showing that the optimally balanced generalization rate can be slower than the minimax prediction benchmark. This result reveals a loss of statistical accuracy in adversarial training. Our analysis shows that this loss arises from the interaction between adversarial robustness and observation noise: the noise contribution in the mixed robustness term slows the approximation rate, although the same term reduces the estimation complexity. To address this limitation, we propose a two-stage noise-debiased procedure that estimates and removes the noise contribution from the mixed term. The resulting estimator improves the generalization rate and attains the minimax polynomial rate, up to a logarithmic factor, when the robustness level is selected at the stated sample-dependent order. Our results characterize the generalization behavior of adversarial training in a nonparametric framework and provide a new interpretation and a principled solution for the trade-off between adversarial robustness and generalization. Numerical experiments support the theoretical findings and demonstrate the effectiveness of the proposed method.