🤖 AI Summary
This paper addresses the sequential release of means and variances over multiple mutually exclusive subsets under user-level differential privacy, assuming heterogeneous data and publicly known per-subset user contribution counts. The goal is to maintain fixed statistical estimation error while mitigating the rapid degradation of privacy loss as the number of subsets increases. We propose an iterative algorithm based on user-contribution suppression and, for the first time, derive exact closed-form expressions for the global sensitivity and worst-case bias of mean and variance estimators under truncation/suppression mechanisms. Theoretically, we prove that this mechanism significantly reduces the cumulative privacy budget consumption rate. Empirically, experiments on both real and synthetic datasets demonstrate that, for a fixed estimation error, the privacy loss degradation factor decreases by several-fold; moreover, when the number of users per subset is fixed, the worst-case estimation error is substantially improved.
📝 Abstract
This paper considers the private release of statistics of several disjoint subsets of a datasets. In particular, we consider the $epsilon$-user-level differentially private release of sample means and variances of sample values in disjoint subsets of a dataset, in a potentially sequential manner. Traditional analysis of the privacy loss under user-level privacy due to the composition of queries to the disjoint subsets necessitates a privacy loss degradation by the total number of disjoint subsets. Our main contribution is an iterative algorithm, based on suppressing user contributions, which seeks to reduce the overall privacy loss degradation under a canonical Laplace mechanism, while not increasing the worst estimation error among the subsets. Important components of this analysis are our exact, analytical characterizations of the sensitivities and the worst-case bias errors of estimators of the sample mean and variance, which are obtained by clipping or suppressing user contributions. We test the performance of our algorithm on real-world and synthetic datasets and demonstrate improvements in the privacy loss degradation factor, for fixed estimation error. We also show improvements in the worst-case error across subsets, via a natural optimization procedure, for fixed numbers of users contributing to each subset.