Score
Designs and implements analytic and numerical methods to compute and bound accumulated privacy loss for differential privacy mechanisms, including composition and amplification analyses (e.g., ODMM composition and subsampling/iteration amplification). Builds precise privacy accountants based on privacy loss distributions (PLDs) and Fourier/FFT convolution methods to evaluate tail probabilities of composed mechanisms (including heterogeneous discrete Gaussian or other mechanism mixes) and to produce rigorous numerical error bounds for the reported privacy guarantees.
This paper addresses privacy accounting for subsampling mechanisms—specifically Poisson and without-replacement sampling—in compositional settings under differential privacy (DP), identifying two prevalent misuses: (i) erroneously assuming the worst-case dataset for a single step suffices for adaptive composition analysis, and (ii) conflating the distinct privacy loss characteristics of the two sampling schemes. Method: We rigorously prove that privacy parameters for subsampled composition cannot be derived by naïvely composing single-step worst-case guarantees. Leveraging Rényi differential privacy and exact privacy loss distribution analysis, we develop a numerical accounting framework incorporating counterexample construction and tight theoretical bounds. Contribution/Results: We establish a decidable criterion for detecting and correcting such misuses, and demonstrate—under typical DP-SGD parameters—that ε values for Poisson and without-replacement sampling may differ by over an order of magnitude. Empirical evaluation confirms our framework prevents significant over- or under-estimation of privacy budgets, substantially improving the reliability of privacy guarantees.
This work addresses the looseness and high computational cost of existing privacy loss analyses under random check-in sampling, which hinder efficient and accurate privacy accounting. The authors propose a general privacy accounting framework based on Privacy Loss Distributions (PLDs), enabling the first mechanism-agnostic and precise privacy analysis for random check-in sampling. This approach overcomes the limitations of traditional methods that rely on approximations or are tailored to specific mechanisms. Theoretical analysis demonstrates that, under the Gaussian mechanism, random check-in sampling achieves a privacy–utility trade-off that is superior or equivalent to Poisson subsampling, making it particularly well-suited for DP-SGD training. Moreover, the proposed framework significantly improves both the efficiency and accuracy of privacy parameter computation.
This work addresses the challenge of privacy parameter computation in differentially private (DP) machine learning under the joint effects of stochastic minibatching and cross-iteration correlated noise—specifically, matrix mechanisms. We propose the first near-exact privacy amplification analysis framework applicable to arbitrary lower-triangular nonnegative correlation matrices. Methodologically, we integrate Monte Carlo privacy accounting with lower-triangular noise modeling, enabling the first tight privacy bound estimation for general correlated noise structures. Our framework supports joint optimization of the correlation matrix under privacy amplification constraints and introduces a practical “ball-and-bin” minibatching mechanism as a robust alternative to Poisson sampling. Experiments demonstrate that our approach achieves lower root-mean-square error (RMSE) than state-of-the-art methods on prefix-sum tasks and significantly improves the privacy–utility trade-off in deep learning tasks.
This paper addresses the sequential release of means and variances over multiple mutually exclusive subsets under user-level differential privacy, assuming heterogeneous data and publicly known per-subset user contribution counts. The goal is to maintain fixed statistical estimation error while mitigating the rapid degradation of privacy loss as the number of subsets increases. We propose an iterative algorithm based on user-contribution suppression and, for the first time, derive exact closed-form expressions for the global sensitivity and worst-case bias of mean and variance estimators under truncation/suppression mechanisms. Theoretically, we prove that this mechanism significantly reduces the cumulative privacy budget consumption rate. Empirically, experiments on both real and synthetic datasets demonstrate that, for a fixed estimation error, the privacy loss degradation factor decreases by several-fold; moreover, when the number of users per subset is fixed, the worst-case estimation error is substantially improved.
This work addresses the composition of interactive differential privacy (DP) mechanisms under *adaptive selection of privacy-loss parameters* and *concurrent invocation*, including interleaved queries and dynamically instantiated mechanisms. To overcome the limitation of existing theory—which fails to model adaptive, concurrent adversarial interaction—the paper establishes, for the first time, that *privacy filters* and *privacy odometers*, originally designed for non-interactive settings, admit rigorous generalization to concurrent interactive settings. This result unifies and proves the robustness of $(varepsilon,delta)$-DP, $f$-DP, and fixed-order Rényi DP under concurrent composition: concurrency itself does not degrade privacy guarantees. The paper develops the first theoretical framework for concurrent interactive DP supporting fully adaptive privacy budget management, and releases an open-source, production-ready implementation—providing both foundational guarantees and practical engineering support for real-world private systems.
This study addresses the complexities of converting between differential privacy notions and the information loss induced by parameter compression. We unify mainstream privacy definitions within an information-theoretic framework by establishing strict equivalence classes among privacy profiles, hypothesis testing, and Rényi Differential Privacy (RDP) curves. This approach precisely delineates conversion boundaries across different privacy concepts and quantifies the accuracy degradation caused by zero-concentrated differential privacy (zCDP) compression. Our results demonstrate that retaining the complete RDP curve reduces noise variance by 45% and improves test accuracy by 8.73 percentage points on the Fashion-MNIST benchmark. Ultimately, this work provides a unified paradigm for the theoretical analysis and optimization of privacy mechanisms.
In the moderate-to-low privacy regime (i.e., small $(\varepsilon, \delta)$), existing Gaussian mechanisms are significantly suboptimal due to excessive noise injection. This work proposes a hybrid Gaussian noise mechanism that constructs a convex combination of multiple Gaussian distributions with identical variances but distinct means, adaptively tuning both the means and mixing weights using sensitivity information. It presents the first systematic construction and analysis of a Gaussian mixture-based perturbation scheme satisfying $(\varepsilon, \delta)$-differential privacy. The authors derive tight variance conditions and an efficient algorithm that substantially reduce both L1 and L2 utility loss in the low-privacy regime, markedly narrowing the performance gap with the theoretically optimal mechanism and achieving near-optimal accuracy.
This work addresses the limitations of existing differential privacy (DP) auditing methods, which predominantly rely on batch sampling and are confined to $(\varepsilon, \delta)$-DP, thereby failing to comprehensively evaluate privacy guarantees under $f$-DP. To overcome this, we propose the first adaptive sequential auditing framework capable of full-spectrum privacy behavior detection for $f$-DP mechanisms without requiring a pre-specified sample size. Our approach operates in both white-box and black-box settings and constructs a sequential hypothesis test grounded in statistical significance theory, leveraging the trade-off function inherent to $f$-DP to identify privacy violations. Theoretical analysis and empirical evaluations demonstrate that our method significantly reduces sampling costs while maintaining statistical power, achieving substantial efficiency gains—particularly in high-overhead scenarios such as DP-SGD—by markedly decreasing the number of required samples.
This work addresses the lack of composability in Pufferfish privacy mechanisms under repeated invocation, which poses a risk of privacy leakage. We establish, for the first time, necessary and sufficient conditions under which Pufferfish mechanisms satisfy linear composability, and construct a formal bridge between Pufferfish and differential privacy that preserves semantic interpretability while enabling strong composability guarantees. Leveraging the $(a,b)$-influence curve framework, we systematically transform existing differential privacy algorithms into composable Pufferfish mechanisms. The resulting algorithms are successfully applied to Markov chain settings, demonstrating significant performance improvements over current approaches.
This work addresses the challenge of precisely characterizing the overall privacy guarantee when composing mechanisms under multiple heterogeneous differential privacy (DP) constraints. The authors propose a general composition framework that, for the first time, enables an exact description of the resulting privacy region after composing an arbitrary number of mechanisms subject to diverse DP bounds. By constructing a binary hypothesis testing–based mixture model and integrating probabilistic mixing with f-DP approximation techniques, the framework yields an exact composition theorem for multiple DP constraints. Moreover, the approach naturally extends to the f-DP setting, significantly enhancing both the tightness and applicability of compositional privacy analysis.