Bounding User Contributions in the Worst-Case for User-Level Differentially Private Mean Estimation

📅 2025-02-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper investigates worst-case error bounds and optimal clipping strategies for mean estimation under user-level ε-differential privacy, focusing on realistic settings with data heterogeneity and heterogeneous user sample sizes. We derive the first exact theoretical characterization of the worst-case error under general clipping mechanisms, yielding the first tight upper bound applicable to non-i.i.d. data and arbitrary per-user sample counts. Building upon this bound, we propose an adaptive optimal clipping mechanism that requires no private estimation of distributional parameters. We prove that our strategy strictly minimizes the worst-case error, thereby overcoming the restrictive assumptions—namely, homogeneous data and fixed sample size—of prior work by Amin et al. (2019). Empirical evaluation demonstrates that our method consistently outperforms state-of-the-art approaches in both i.i.d. and non-i.i.d. settings.

Technology Category

Machine Learning: PrivacyReasoning under Uncertainty: Stochastic OptimizationSearch and Optimization: Mixed Discrete/Continuous Search

Application Category

User Modeling, Personalization and Recommendation: User privacy protection in personalized systemsSecurity and Privacy: Large-scale security measurementsResponsible Web: Data and user privacy-enhancing technologies for the Web
📝 Abstract
In this article, we revisit the well-studied problem of mean estimation under user-level $varepsilon$-differential privacy (DP). While user-level $varepsilon$-DP mechanisms for mean estimation, which typically bound (or clip) user contributions to reduce sensitivity, are well-known, an analysis of their estimation errors usually assumes that the data samples are independent and identically distributed (i.i.d.), and sometimes also that all participating users contribute the same number of samples (data homogeneity). Our main result is a precise characterization of the emph{worst-case} estimation error under general clipping strategies, for heterogeneous data, and as a by-product, the clipping strategy that gives rise to the smallest worst-case error. Interestingly, we show via experimental studies that even for i.i.d. samples, our clipping strategy performs uniformly better that the well-known clipping strategy of Amin et al. (2019), which involves additional, private parameter estimation.
Problem

Research questions and friction points this paper is trying to address.

Worst-case error in mean estimation
User-level differential privacy mechanisms
Clipping strategies for heterogeneous data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Worst-case error analysis
Heterogeneous data handling
Optimal clipping strategy