🤖 AI Summary
This study addresses the challenge of quantifying differential privacy budgets to balance model utility against data reconstruction risks. By deriving entropy lower bounds, analyzing output perturbations in linear regression, and employing zero-concentrated differential privacy modeling, this work reveals a sharp phase transition between privacy budgets and data dimensionality, refining threshold determination from the full-dimensional space to the effective subspace dimension. The authors rigorously delineate the information-theoretically unrecoverable region from the practically feasible attack domain. These theoretical predictions are empirically validated on synthetic datasets as well as CIFAR-10 and ImageNet. Ultimately, this research establishes a solid theoretical foundation for the principled selection of privacy parameters.
📝 Abstract
Data reconstruction attacks have empirically been successful in recovering training samples from learned models, raising privacy concerns and motivating defenses with guarantees that remain valid against future threats. While differential privacy (DP) provides formal protection, choosing the privacy budget remains a challenge: small budgets severely reduce utility, but it is hard to quantify how large the budget can be without allowing accurate reconstruction. In this work, we study informed attackers who aim to reconstruct a single $d$-dimensional training sample from a $ρ$-zero-concentrated DP model, knowing all other training data. Our main contribution is to establish a sharp transition at $ρ\asymp d$ for data reconstruction: on the one hand, we derive entropy-based lower bounds for any private mechanism and any attack, characterizing a set of target priors for which reconstruction is information-theoretically impossible for $ρ\ll d$; on the other hand, we analyze a simple attack on private linear regression with output perturbation, showing that reconstruction is practically feasible for $ρ\gg d$. Remarkably, the transition moves to $ρ\asymp s$ for data lying in an $s$-dimensional subspace, demonstrating that the privacy budget guaranteeing adequate protection must be assessed in terms of the effective dimension of the data. We validate our findings via experiments on synthetic data and natural images (CIFAR-10, ImageNet).