🤖 AI Summary
This study addresses the reliance of DP-SGD privacy budget calibration on computationally expensive numerical searches and the absence of efficient closed-form solutions. Focusing on randomly assigned DP-SGD, this work leverages chi-square divergence and Gaussian mixture model theory to derive a closed-form upper bound on the accuracy of membership inference attacks, yielding the first concise analytical bound applicable to this setting. The proposed approach enables microsecond-level noise calibration while reducing the noise overhead by approximately 50% compared to existing state-of-the-art closed-form bounds. Experimental results confirm that empirical attack accuracy remains strictly below the theoretical bound, demonstrating that the method effectively balances computational efficiency with rigorous privacy guarantees.
📝 Abstract
DP-SGD protects training data by adding Gaussian noise to clipped gradients. The amount of noise is usually chosen by running a numerical privacy accountant inside a search. We study DP-SGD with random allocation, where each epoch uses every record once, at a randomly chosen step. For this setting we give a one-line formula that bounds the accuracy of every membership inference attack (MIA) on the trained model. With $M$ steps per epoch, $E$ epochs and noise multiplier $σ$, and with membership and non-membership equally likely a priori, the attack accuracy is at most $\frac12+\frac14\sqrt{(1+(e^{1/σ^2}-1)/M)^E-1}$. The formula comes from the chi-square divergence between a Gaussian distribution and a Gaussian mixture that dominates random allocation. It is interpretable and gives $σ$ in about a microsecond. Where applicable, our formula needs at most about half the noise of the state-of-the-art closed-form bound. To measure how close the bound is, we also derive an exact expression for the attack accuracy of these two distributions and evaluate it numerically. Calibrating to this exact expression requires $13.0\%$ to $20.2\%$ less noise than the formula in our main experiments, and since it is exact, no accountant that knows only $M$, $E$ and $σ$ can certify a smaller $σ$. In training, the resulting $σ$ outperforms the formula and matches a published accountant in test accuracy. It is found in seconds and certified in minutes, whereas every search we ran with that accountant took longer or returned at least $0.62\%$ more noise. We show that MIAs on the trained models stay below the bound.