private dataset embedding

Design and implement methods that compute a single low-dimensional numerical representation of an entire dataset by embedding and aggregating example-level features, then apply differential privacy mechanisms (e.g., calibrated noise added to the aggregated embedding) to produce a privatized dataset representation with formal privacy guarantees; analyze and tune the privacy–utility tradeoffs, aggregation strategies, and any post-processing (dimensionality reduction, normalization) required to use the private embedding.

privatedatasetembedding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Power Mechanism: Private Tabular Representation Release for Model Agnostic Consumption

Oct 07, 2025
PV
Praneeth Vepakomma
🏛️ MBZUAI | MIT | EPFL

This work addresses privacy risks arising from sharing data embeddings (activations) in collaborative learning—an underexplored setting compared to weight-sharing. Existing approaches lack formal privacy guarantees for embedding sharing and struggle to accommodate heterogeneous server-side models. We propose the first differentially private mechanism specifically designed for embedding sharing. Our method jointly designs a privacy-preserving encoder network and a lightweight utility generation network, enabling high-accuracy, low-overhead private embedding generation in a single communication round. Crucially, the mechanism is model-agnostic: it imposes no structural assumptions on the server-side model and seamlessly integrates with diverse downstream models—including deep neural networks, random forests, and XGBoost. Experiments demonstrate that, under strict (ε,δ)-differential privacy, our approach significantly reduces client-side computational overhead while maintaining performance close to non-private baselines across multiple tasks and model architectures.

Developing differentially private embedding sharing mechanismsEnabling model-agnostic consumption of privatized tabular dataReducing client computation and communication rounds

Certification for Differentially Private Prediction in Gradient-Based Training

Jun 19, 2024
MW
Matthew Wicker
🏛️ Imperial College London | The Alan Turing Institute | Accenture Labs | ETH Zurich | LogicStar.ai | University of Cambridge

Differential privacy (DP) gradient training suffers from excessive noise injection and suboptimal privacy–utility trade-offs due to reliance on global sensitivity, which is overly conservative for modern deep models. Method: This paper proposes the first scalable and verifiable framework for computing upper bounds on both local and smooth sensitivity—novelly integrating convex relaxation with interval-bound propagation to enable precise, efficient estimation of smooth sensitivity during gradient computation in contemporary deep neural networks. Contribution/Results: Our approach overcomes longstanding theoretical and computational barriers in rigorously bounding sensitivity. Experiments across financial risk assessment, medical image classification, and multi-task NLP demonstrate that our method reduces required noise magnitude by an order of magnitude, yielding substantial improvements in prediction accuracy and practical utility under identical privacy budgets (e.g., ε = 2, δ = 10⁻⁵). The framework provides stronger theoretical guarantees for private inference while ensuring engineering feasibility and scalability.

Achieving differential privacy in prediction via noise additionEnhancing private prediction accuracy in medical and NLP tasksImproving privacy-utility trade-offs with dataset-specific sensitivity bounds

Private Training&Data Generation by Clustering Embeddings

Jun 20, 2025
FZ
Felix Zhou
🏛️ Yale University | Texas A&M University | Google Research

To address the privacy risk wherein deep learning models inadvertently memorize and leak raw samples in sensitive-data scenarios, this paper proposes a differential privacy (DP)-driven synthetic embedding generation method. Our approach jointly integrates DP clustering with a Gaussian Mixture Model (GMM) in the embedding space to generate high-fidelity synthetic data with provable privacy guarantees. Notably, we are the first to incorporate DP clustering theory into the GMM fitting process, enabling flexible substitution of encoders and decoders while preserving generality and scalability. Experiments on standard benchmarks demonstrate that training lightweight two-layer networks on our synthetic embeddings achieves state-of-the-art (SOTA) classification accuracy. Moreover, images synthesized via our method attain downstream task performance comparable to current best approaches.

Achieving high classification accuracy with synthetic embeddingsGenerating synthetic datasets to preserve differential privacyPrivate training of deep neural networks with sensitive data

This study addresses a fundamental tension between differential privacy and data valuation: the former requires insensitivity to individual records, while the latter demands precise quantification of each data point’s contribution. The work systematically analyzes how mainstream valuation methods—such as Shapley values and influence functions—fail under differential privacy constraints, identifying high-sensitivity components within these algorithms. It proposes design principles for privacy-friendly valuation mechanisms and employs sensitivity analysis alongside privacy-utility trade-off evaluations to reveal the limitations of current approaches in preserving the discriminative power of rare samples. By delineating the feasible boundaries of private data valuation, this research lays a theoretical foundation for developing practical mechanisms that jointly uphold privacy guarantees and valuation utility.

data valuationdifferential privacyheterogeneous data

Training Differentially Private Models with Secure Multiparty Computation

Feb 05, 2022
SP
Sikha Pentyala
🏛️ University of Washington Tacoma | Mila – Quebec AI Institute | University of Brasilia | Monash University | Ghent University

Balancing privacy preservation and model accuracy remains challenging in collaborative modeling among multiple data owners. Method: This paper proposes a novel framework that deeply integrates differential privacy (DP) with secure multi-party computation (MPC). It is the first to provably inject Laplacian noise directly within an MPC protocol—performing privacy-parameter perturbation under secret sharing during distributed gradient computation. This ensures strict ε-differential privacy guarantees while avoiding the accuracy degradation typically caused by global noise in conventional DP approaches. Contribution/Results: The method enables privacy-preserving joint training on highly sensitive data (e.g., genomic data) without exposing raw samples. It achieved first place in the iDASH 2021 Track III competition, significantly outperforming pure-DP baselines in accuracy. By unifying formal privacy guarantees with practical efficiency, this work establishes a new paradigm for privacy-enhancing technologies that simultaneously satisfies rigorous security requirements and real-world usability.

Combining MPC and DP to improve accuracy without privacy lossEnsuring formal privacy guarantees for each owner's dataLearning machine learning models from multiple data owners

Latest Papers

What's happening recently
View more

This study addresses the privacy-utility imbalance caused by global sensitivity in differential privacy and the absence of finite privacy guarantees for unbounded regression. It pioneers the extension of abstract interpretation to private prediction for continuous unbounded regression. By proposing an Abstract Gradient Sampling (AGS) algorithm alongside a smooth sensitivity upper-bounding technique, this work reformulates parameter learning as a regression problem, enabling formal certification of private learning via reachability analysis. Experimental results demonstrate that the derived regression bounds are tighter than those obtained using global sensitivity baselines. Furthermore, the proposed approach achieves the first finite privacy guarantees in unbounded settings and yields private learning performance superior to standard algorithms under matched conditions.

Differential PrivacyFormal MethodsPrivate Learning

This work addresses the unnecessary utility loss in existing differential privacy (DP) training methods, which add noise to both inputs and labels even when labels are public. To remedy this, the authors propose an end-to-end DP framework that protects only the input data. Their approach introduces the Dirichlet mechanism—applied for the first time to the softmax output layer—to enforce input-level privacy, and leverages Rényi differential privacy theory to tightly track the privacy budget across multiple training epochs. Evaluated on standard benchmarks such as CIFAR-10 and MNIST, the method achieves state-of-the-art accuracy under the label-public setting: on CIFAR-10, it attains 88.17% accuracy at ε=4—nearly a 10-percentage-point improvement over prior art—and maintains 82.96% accuracy even at the stringent privacy level of ε=1, setting a new performance record for DP training with public labels.

data privacydeep neural networksdifferential privacy

This work addresses the challenge of dynamically satisfying varying differential privacy (DP) requirements during inference without retraining models. To this end, it proposes two training-free post-processing methods—random selection and linear combination—that generate new models meeting arbitrary target DP parameters by fusing pre-trained models representing different privacy-utility trade-offs. As the first systematic study to leverage model fusion for adaptively fulfilling arbitrary DP guarantees, the paper provides rigorous theoretical analysis grounded in Rényi differential privacy and privacy loss distributions. It further proves that linear combination strictly dominates random selection in terms of the privacy-utility trade-off. Empirical evaluations on both synthetic and real-world datasets validate the effectiveness and practicality of the proposed approaches.

Differential PrivacyModel MergingPost Processing

This study addresses key challenges in privacy-preserving transfer learning, where evaluating the utility of external data solely through summary statistics, mitigating negative transfer, and balancing privacy noise against predictive performance remain difficult. To this end, the authors develop an error theory grounded in high-dimensional asymptotic analysis and weighted ridge regression that captures the interplay among sample size, covariance structure, model shift, and privacy noise. By integrating zero-concentrated differential privacy (zCDP) with certainty equivalence techniques, they construct a test error expression from aggregated statistics to guide hyperparameter optimization. This framework enables utility assessment and decision support without requiring access to raw data. Extensive experiments on both synthetic and real-world datasets validate the theoretical findings, establishing a computationally tractable foundation for data selection under strict privacy constraints.

Dataset SelectionDifferential PrivacyHigh-Dimensional Asymptotics

This work addresses a central challenge in differential privacy: enhancing algorithmic utility without compromising privacy guarantees or incurring excessive computational complexity. The authors propose a post-processing denoising method grounded in empirical Bayes estimation, which effectively reduces mean squared error using only the outputs of Gaussian differential privacy mechanisms. To the best of our knowledge, this is the first systematic application of the empirical Bayes framework to differential privacy post-processing. The approach significantly improves utility without altering the underlying privacy mechanism, offering both simplicity and broad applicability. Empirical evaluations demonstrate consistent performance gains over existing differentially private algorithms across diverse tasks, including histogram release, principal component analysis, and linear regression.

differential privacyempirical BayesGaussian mechanism

Hot Scholars

MG

Matthias Guckenberger

Professor for Radiation Oncology, University Hospital Zurich, University of Zurich
YX

Yuanyuan Xu

University of New South Wales
Graph Neural NetworksBig Data
CB

Christoph Bert

Professor für Medizinische Strahlenphysik, FAU Erlangen-Nürnberg
radiation oncologymedical physics
MI

Mohsen Imani

Associate Professor, University of California Irvine
Machine LearningBrain-Inspired SystemsIntelligent Systems
RF

Reza Fotohi

Post-Doctoral Researcher at Shahid Beheshti University
SecurityPrivacyAI/MLLLMs