Score
Design and analyze stochastic encoders that convert outputs of privacy mechanisms into Poisson-based representations (randomized codes whose components are modeled as Poisson counts), producing compact, adjustable-rate encodings that preserve differential privacy guarantees and allow tractable privacy accounting after encoding.
This work addresses the lack of efficient compression mechanisms for high-dimensional data, such as images, under differential privacy, which leads to substantial storage overhead and limited practicality. The authors propose DP-DiPP, a novel framework that uniquely integrates Poisson Private Representations (PPR) with the diffusion-based compression method DiffC, leveraging stochastic encoding and diffusion models to achieve a flexible trade-off among privacy, compression ratio, and utility. Evaluated on private image classification using CIFAR-10, DP-DiPP achieves 10–30 times higher compression ratios compared to existing baselines while maintaining comparable privacy guarantees and model utility.
This work addresses the looseness and high computational cost of existing privacy loss analyses under random check-in sampling, which hinder efficient and accurate privacy accounting. The authors propose a general privacy accounting framework based on Privacy Loss Distributions (PLDs), enabling the first mechanism-agnostic and precise privacy analysis for random check-in sampling. This approach overcomes the limitations of traditional methods that rely on approximations or are tailored to specific mechanisms. Theoretical analysis demonstrates that, under the Gaussian mechanism, random check-in sampling achieves a privacy–utility trade-off that is superior or equivalent to Poisson subsampling, making it particularly well-suited for DP-SGD training. Moreover, the proposed framework significantly improves both the efficiency and accuracy of privacy parameter computation.
This paper addresses privacy accounting for subsampling mechanisms—specifically Poisson and without-replacement sampling—in compositional settings under differential privacy (DP), identifying two prevalent misuses: (i) erroneously assuming the worst-case dataset for a single step suffices for adaptive composition analysis, and (ii) conflating the distinct privacy loss characteristics of the two sampling schemes. Method: We rigorously prove that privacy parameters for subsampled composition cannot be derived by naïvely composing single-step worst-case guarantees. Leveraging Rényi differential privacy and exact privacy loss distribution analysis, we develop a numerical accounting framework incorporating counterexample construction and tight theoretical bounds. Contribution/Results: We establish a decidable criterion for detecting and correcting such misuses, and demonstrate—under typical DP-SGD parameters—that ε values for Poisson and without-replacement sampling may differ by over an order of magnitude. Empirical evaluation confirms our framework prevents significant over- or under-estimation of privacy budgets, substantially improving the reliability of privacy guarantees.
This work reveals that replacing Poisson subsampling with data shuffling in DP-SGD leads to severe overestimation of theoretical differential privacy guarantees. To address this, the authors introduce the first empirical privacy auditing framework specifically designed for DP-SGD with shuffling, integrating differential privacy auditing, membership inference attacks, and empirical leakage measurement. They conduct a systematic evaluation across two realistic shuffling variants. Experiments show that mainstream implementations overestimate the privacy budget ε by 2–4× on average, resulting in substantially increased privacy leakage; the degree of overestimation is significantly influenced by batch size, prescribed ε, and threat model. This study provides the first quantitative characterization of the privacy gap induced by shuffling, delivering critical warnings for secure DP-SGD deployment and establishing a reproducible benchmark for empirical privacy assessment.
This work addresses the challenge of privacy parameter computation in differentially private (DP) machine learning under the joint effects of stochastic minibatching and cross-iteration correlated noise—specifically, matrix mechanisms. We propose the first near-exact privacy amplification analysis framework applicable to arbitrary lower-triangular nonnegative correlation matrices. Methodologically, we integrate Monte Carlo privacy accounting with lower-triangular noise modeling, enabling the first tight privacy bound estimation for general correlated noise structures. Our framework supports joint optimization of the correlation matrix under privacy amplification constraints and introduces a practical “ball-and-bin” minibatching mechanism as a robust alternative to Poisson sampling. Experiments demonstrate that our approach achieves lower root-mean-square error (RMSE) than state-of-the-art methods on prefix-sum tasks and significantly improves the privacy–utility trade-off in deep learning tasks.
本文探讨了在不独立情况下通过负相关参与机制保证隐私放大,对比泊松子采样方法,明确了其在高斯机制下的适用性和限制。
本文针对实例编码隐私保护方法的可逆性问题,提出一种新的谱感知界限方法,该方法更紧致、适用于确定性编码器并扩展到多种相似度量。
This work transcends the limitations of traditional scalar privacy loss by adopting a geometric perspective to precisely characterize the random triangular structures induced by high-dimensional noise perturbations in differential privacy. By constructing two complementary geometric representations—mapping the random triangle formed by sensitivity and noise vectors onto a simplex and a hemisphere—it establishes, for the first time, an exact geometric connection between differential privacy and probabilistic shape analysis. Through rigorous derivations of probability densities, geometric mappings, spectral shape analysis, and high-dimensional asymptotic methods, the study reformulates classical privacy loss variables, revealing an elliptical support structure on the simplex and phenomena of equatorial drift and zonal concentration on the hemisphere. These insights offer a novel geometric framework for designing privacy-preserving mechanisms.
This work addresses critical limitations of the conventional discrete Gaussian mechanism in differential privacy, which is vulnerable to floating-point precision issues and demands substantial high-quality randomness. The authors propose the dithered Gaussian mechanism, which decouples randomness into a privacy-critical high-quality component and a non-critical, computationally efficient component through output-side discretization and dual-source randomization. This approach preserves the theoretical privacy guarantees of the standard Gaussian mechanism while eliminating floating-point security vulnerabilities. Notably, it drastically reduces the requirement for high-quality random bits—rendering this demand independent of noise magnitude—and enables cryptographically secure noise generation in DP-SGD with minimal computational overhead, thereby achieving a strong balance between security and practicality.
This study addresses the complexities of converting between differential privacy notions and the information loss induced by parameter compression. We unify mainstream privacy definitions within an information-theoretic framework by establishing strict equivalence classes among privacy profiles, hypothesis testing, and Rényi Differential Privacy (RDP) curves. This approach precisely delineates conversion boundaries across different privacy concepts and quantifies the accuracy degradation caused by zero-concentrated differential privacy (zCDP) compression. Our results demonstrate that retaining the complete RDP curve reduces noise variance by 45% and improves test accuracy by 8.73 percentage points on the Fashion-MNIST benchmark. Ultimately, this work provides a unified paradigm for the theoretical analysis and optimization of privacy mechanisms.