Adapting to noise tails in private linear regression

📅 2026-03-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses robust linear regression under heavy-tailed noise while preserving differential privacy. The authors propose an estimation framework that integrates a tunable Huber loss with a differentially private mechanism. In low-dimensional settings, they employ noisy truncated gradient descent, and in high-dimensional sparse regimes, they adopt a noisy iterative hard thresholding algorithm, achieving a balance among privacy, robustness, and statistical efficiency. Theoretical analysis precisely characterizes the non-asymptotic convergence rates, revealing their dependence on the moment exponent, privacy parameters, sample size, and intrinsic dimensionality, and elucidating the trade-offs among bias, privacy, and robustness. The method attains near-optimal rates under sub-Gaussian errors and demonstrates strong empirical performance under heavy-tailed noise, as validated by both theoretical guarantees and experiments on synthetic data and two real-world datasets.

Technology Category

Machine Learning: PrivacyNatural Language Processing: Safety and RobustnessIntelligent Robots: State Estimation

Application Category

Security and Privacy: Large-scale security measurementsGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methods
📝 Abstract
While the traditional goal of statistics is to infer population parameters, modern practice increasingly demands protection of individual privacy. One way to address this need is to adapt classical statistical procedures into privacy-preserving algorithms. In this paper, we develop differentially private tail-robust methods for linear regression. The trade-off among bias, privacy, and robustness is controlled by a tunable robustification parameter in the Huber loss. We implement noisy clipped gradient descent for low-dimensional settings and noisy iterative hard thresholding for high-dimensional sparse models. Under sub-Gaussian errors, our method achieves near-optimal convergence rates while relaxing several assumptions required in earlier work. For heavy-tailed errors, we explicitly characterize how the non-asymptotic convergence rate depends on the moment index, privacy parameters, sample size, and intrinsic dimension. Our analysis shows how the moment index influences the choice of robustification parameters and, in turn, the resulting statistical error and privacy cost. By quantifying the interplay among bias, privacy, and robustness, we extend classical perspectives on privacy-preserving robust regression. The proposed methods are evaluated through simulations and two real datasets.
Problem

Research questions and friction points this paper is trying to address.

differential privacy
robust regression
heavy-tailed errors
linear regression
bias-privacy-robustness trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

differentially private regression
heavy-tailed noise
Huber loss
robustness-privacy trade-off
iterative hard thresholding
🔎 Similar Papers
J
Jinyuan Chang
Joint Laboratory of Data Science and Business Intelligence, Institute of Statistical Interdisciplinary Research, Southwestern University of Finance and Economics, Chengdu, China
L
Lin Yang
Joint Laboratory of Data Science and Business Intelligence, Institute of Statistical Interdisciplinary Research, Southwestern University of Finance and Economics, Chengdu, China
M
Mengyue Zha
Department of Mathematics, Hong Kong University of Science and Technology, Hong Kong
W
Wen-Xin Zhou
College of Business Administration, University of Illinois Chicago, Chicago, IL, USA