🤖 AI Summary
This paper investigates worst-case error bounds and optimal clipping strategies for mean estimation under user-level ε-differential privacy, focusing on realistic settings with data heterogeneity and heterogeneous user sample sizes. We derive the first exact theoretical characterization of the worst-case error under general clipping mechanisms, yielding the first tight upper bound applicable to non-i.i.d. data and arbitrary per-user sample counts. Building upon this bound, we propose an adaptive optimal clipping mechanism that requires no private estimation of distributional parameters. We prove that our strategy strictly minimizes the worst-case error, thereby overcoming the restrictive assumptions—namely, homogeneous data and fixed sample size—of prior work by Amin et al. (2019). Empirical evaluation demonstrates that our method consistently outperforms state-of-the-art approaches in both i.i.d. and non-i.i.d. settings.
📝 Abstract
In this article, we revisit the well-studied problem of mean estimation under user-level $varepsilon$-differential privacy (DP). While user-level $varepsilon$-DP mechanisms for mean estimation, which typically bound (or clip) user contributions to reduce sensitivity, are well-known, an analysis of their estimation errors usually assumes that the data samples are independent and identically distributed (i.i.d.), and sometimes also that all participating users contribute the same number of samples (data homogeneity). Our main result is a precise characterization of the emph{worst-case} estimation error under general clipping strategies, for heterogeneous data, and as a by-product, the clipping strategy that gives rise to the smallest worst-case error. Interestingly, we show via experimental studies that even for i.i.d. samples, our clipping strategy performs uniformly better that the well-known clipping strategy of Amin et al. (2019), which involves additional, private parameter estimation.