🤖 AI Summary
This work addresses robust linear regression under heavy-tailed noise while preserving differential privacy. The authors propose an estimation framework that integrates a tunable Huber loss with a differentially private mechanism. In low-dimensional settings, they employ noisy truncated gradient descent, and in high-dimensional sparse regimes, they adopt a noisy iterative hard thresholding algorithm, achieving a balance among privacy, robustness, and statistical efficiency. Theoretical analysis precisely characterizes the non-asymptotic convergence rates, revealing their dependence on the moment exponent, privacy parameters, sample size, and intrinsic dimensionality, and elucidating the trade-offs among bias, privacy, and robustness. The method attains near-optimal rates under sub-Gaussian errors and demonstrates strong empirical performance under heavy-tailed noise, as validated by both theoretical guarantees and experiments on synthetic data and two real-world datasets.
📝 Abstract
While the traditional goal of statistics is to infer population parameters, modern practice increasingly demands protection of individual privacy. One way to address this need is to adapt classical statistical procedures into privacy-preserving algorithms. In this paper, we develop differentially private tail-robust methods for linear regression. The trade-off among bias, privacy, and robustness is controlled by a tunable robustification parameter in the Huber loss. We implement noisy clipped gradient descent for low-dimensional settings and noisy iterative hard thresholding for high-dimensional sparse models. Under sub-Gaussian errors, our method achieves near-optimal convergence rates while relaxing several assumptions required in earlier work. For heavy-tailed errors, we explicitly characterize how the non-asymptotic convergence rate depends on the moment index, privacy parameters, sample size, and intrinsic dimension. Our analysis shows how the moment index influences the choice of robustification parameters and, in turn, the resulting statistical error and privacy cost. By quantifying the interplay among bias, privacy, and robustness, we extend classical perspectives on privacy-preserving robust regression. The proposed methods are evaluated through simulations and two real datasets.