differentially private sampling

Designs, implements, and evaluates probabilistic sampling mechanisms that produce outputs under formal differential privacy guarantees, including variants of the exponential mechanism (e.g., mew sampling), marked edge walk samplers, and private plan-sampling procedures. Builds and analyzes algorithms that target score or utility functions while calibrating privacy budgets and acceptance probabilities, quantifying privacy–utility tradeoffs and robustness to adversarial inputs.

differentiallyprivatesampling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Avoiding Pitfalls for Privacy Accounting of Subsampled Mechanisms under Composition

May 27, 2024
CL
C. Lebeda
🏛️ IT University of Copenhagen | University of Waterloo | Vector Institute | Google DeepMind

This paper addresses privacy accounting for subsampling mechanisms—specifically Poisson and without-replacement sampling—in compositional settings under differential privacy (DP), identifying two prevalent misuses: (i) erroneously assuming the worst-case dataset for a single step suffices for adaptive composition analysis, and (ii) conflating the distinct privacy loss characteristics of the two sampling schemes. Method: We rigorously prove that privacy parameters for subsampled composition cannot be derived by naïvely composing single-step worst-case guarantees. Leveraging Rényi differential privacy and exact privacy loss distribution analysis, we develop a numerical accounting framework incorporating counterexample construction and tight theoretical bounds. Contribution/Results: We establish a decidable criterion for detecting and correcting such misuses, and demonstrate—under typical DP-SGD parameters—that ε values for Poisson and without-replacement sampling may differ by over an order of magnitude. Empirical evaluation confirms our framework prevents significant over- or under-estimation of privacy budgets, substantially improving the reliability of privacy guarantees.

Clarifying misconceptions about worst-case dataset assumptions in compositionComparing privacy differences between Poisson and without-replacement samplingComputing tight privacy guarantees for composed subsampled mechanisms

Privacy Mechanism Design based on Empirical Distributions

Sep 26, 2025
LG
Leonhard Grosse
🏛️ KTH Royal Institute of Technology | Inria Saclay

This paper addresses the challenge of achieving pointwise maximal leakage (PML) privacy guarantees when the underlying data-generating distribution is unknown. We propose the first distribution-agnostic framework for designing PML-private mechanisms. Methodologically, we extend PML to distributional uncertainty sets, derive robust (ε,δ)-PML guarantees via empirical distributions and large-deviation theory, and formulate mechanism design as a tractable convex optimization problem with linear constraints. Theoretically, our framework provides rigorous, verifiable privacy guarantees; empirically, on binary data, it achieves significantly higher utility than local differential privacy under identical privacy budgets, while remaining compatible with Laplace and Gaussian mechanisms. Our core contribution lies in bridging the theoretical gap between empirical distribution estimation and PML privacy design—establishing a novel privacy paradigm that simultaneously ensures statistical robustness and computational feasibility.

Achieving utility improvements while maintaining privacy guaranteesDesigning privacy mechanisms using empirical data distributionsProviding worst-case leakage bounds for uncertain distributions

Beyond the Calibration Point: Mechanism Comparison in Differential Privacy

Jun 13, 2024
GK
G. Kaissis
🏛️ Technical University of Munich | LMU Munich | Google DeepMind

Differential privacy (DP) mechanisms are commonly reported at a single $(varepsilon,delta)$ point, obscuring substantial differences in actual privacy risk among mechanisms sharing identical $(varepsilon,delta)$ parameters—leading to systematic underestimation of risk. Method: We propose a unified quantification framework grounded in $Delta$-divergence, integrating f-differential privacy, Bayesian privacy interpretations, and Blackwell order theory for the first time to establish a decision-theoretically principled paradigm for comparing DP mechanisms. Contribution/Results: By rigorously characterizing worst-case privacy vulnerability disparities, we expose non-negligible excess risk in mainstream noise mechanisms used in DP-SGD. Our framework yields a verifiable, ordinal privacy strength assessment tool—enabling rigorous, theoretically grounded selection of privacy-preserving mechanisms.

Addressing gaps in current DP-SGD privacy risk understandingComparing DP mechanisms beyond single (ε, δ) pairsQuantifying worst-case excess privacy vulnerabilities

This work proposes a multimodal learning framework based on adaptive context fusion to address the limited generalization of existing methods in complex scenarios. By dynamically aligning visual and linguistic features and incorporating a lightweight gating mechanism, the approach enables efficient cross-modal information integration. Experimental results demonstrate that the model significantly outperforms current state-of-the-art methods across multiple benchmark datasets, exhibiting notably enhanced robustness under low-resource and noisy conditions. The primary contribution lies in the design of a scalable fusion architecture that offers a novel perspective for multimodal representation learning while achieving a favorable balance between computational efficiency and performance.

Collusion AttackExcess VulnerabilityIndividual Differential Privacy

Existing local differential privacy (LDP) mechanisms for numeric data lack a unified optimization framework for arbitrary finite output cardinality (N); optimal perturbation schemes are known only for degenerate cases where the output space size (|mathcal{Y}|) is either extremely small ((2) or (3)) or infinite. Method: We propose the first general-purpose LDP mechanism adaptable to any discrete output size (N), derived by jointly optimizing the minimum-variance unbiased estimation problem under LDP constraints. Our approach integrates closed-form analytical derivation with efficient numerical optimization and naturally extends to mean, variance, and distribution estimation. Contribution/Results: The mechanism achieves Pareto-optimal trade-offs between estimation accuracy and privacy. Experiments demonstrate state-of-the-art accuracy across multiple statistical estimation tasks, with low communication overhead and significant improvements over existing LDP baselines.

Extending LDP framework for accurate distribution estimation tasksMinimizing estimation variance in numerical data collectionOptimizing LDP mechanisms for arbitrary discrete output sizes

Latest Papers

What's happening recently
View more

This work addresses the design of optimal mechanisms for binary hypothesis testing under ε-local differential privacy (LDP). It proposes the Sort-Partition-Randomize (SPR) framework, which first orders input symbols by their likelihood ratios, partitions them into contiguous blocks, and then applies randomized response to the block labels. Leveraging this structure, the paper establishes the existence of an optimal mechanism for any privacy budget ε and any f-divergence–based utility objective—including total variation distance and KL divergence—and presents, for the first time, a dynamic programming algorithm that computes such a mechanism exactly in O(k³) time. This approach overcomes prior limitations restricted to asymptotic privacy regimes, enabling efficient computation of optimal mechanisms across the full range of privacy parameters.

binary hypothesis testingf-divergencelocal differential privacy

This study addresses periodic artifacts in the OpenDP discrete Laplace sampler caused by underlying library defects, which compromise the statistical reliability of differential privacy mechanisms. To resolve this, we propose a fault diagnosis strategy that isolates nested sampling layers to precisely identify failing components, and reconstructs Bernoulli primitives using exact rational arithmetic to eliminate artifacts at their source. Statistical validation over one million samples demonstrates that the corrected sampler output exhibits no significant deviation from the theoretical distribution. This work effectively resolves sampling bias in the discrete Laplace distribution within differential privacy frameworks, enhancing the robustness of privacy-preserving algorithms in safety-critical scenarios.

differential privacydiscrete Laplace samplerOpenDP

This work addresses the problem of efficiently generating synthetic data under differential privacy for a given family of queries. By parameterizing the problem with the treewidth of the query family’s associated graph, the authors establish—for the first time—that the problem is fixed-parameter tractable. They propose a unified dynamic programming framework that integrates linear programming duality-based separation, subsampled private multiplicative weights, and Gibbs sampling techniques. This approach achieves theoretically optimal error rates across the full parameter regime, significantly enhancing both the scalability and practical utility of differentially private synthetic data generation.

differential privacyfixed-parameter tractabilityincidence graph

This work investigates privacy leakage arising from releasing posterior sample paths of Gaussian processes under the strict setting where training data are entirely private. It establishes, for the first time, that the inherent randomness of posterior sampling naturally provides differential privacy guarantees, and derives rigorous privacy bounds using Rényi differential privacy theory. The study proposes effective ridge regularization as a core mechanism to control privacy levels, complemented by calibrated noise injection for enhanced protection. Both theoretical analysis and empirical results demonstrate that the degree of privacy leakage is significantly influenced by the strength of regularization, posterior variance, and the number of released samples. In settings with noisy observations, moderate regularization effectively safeguards privacy while preserving utility for downstream tasks.

Differential PrivacyGaussian ProcessPosterior Sampling

Hot Scholars

AS

Anthony Simonet-Boulogne

Head of Research & Innovation at iExec Blockchain Tech
Decentralized and Distributed ComputingDistributed Data ManagementEdge ComputingCloud Computing
YB

Yacine Belal

CEA LIST
Distributed SystemsMachine LearningPrivacy Enhancing Technologies
HF

Hendrik Fichtenberger

Google Research
clusteringgraph algorithmsproperty testingsublinear algorithms
SB

Sonia Ben Mokhtar

LIRIS CNRS
Distributed systemsFault tolerancePrivacyDistributed Machine Learning
MH

Monika Henzinger

Professor of Computer Science, Institute of Science and Technology Austria (ISTA)
Efficient combinatorial algorithms