DPDSyn: Improving Differentially Private Dataset Synthesis for Model Training by Downstream Task Guidance

📅 2026-04-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of generating highly useful synthetic data under differential privacy constraints, particularly when modeling low-dimensional distributions proves difficult and leads to utility degradation. The authors propose a novel approach that first trains a differentially private model for the target downstream task on the original data and then leverages this model to guide the synthesis process—marking the first use of a differentially private model’s task-specific capability to inform data generation. The method substantially enhances both the quality and efficiency of synthetic data. Evaluated across four benchmark datasets against eight state-of-the-art baselines, it achieves up to a 2.40× improvement in downstream task accuracy and a 333.73× gain in generation speed, while demonstrating strong scalability.

Technology Category

Machine Learning: PrivacyNatural Language Processing: GenerationComputer Vision: Diffusion Models for Vision

Application Category

User Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSecurity and Privacy: Data transparency and provenanceEconomics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAI
📝 Abstract
How to synthesize a dataset while achieving differential privacy for AI model training is a meaningful but challenging problem. To address this problem, state-of-the-art methods first select original private dataset's multiple low-dimensional distributions that have the potential to approximate the distribution of original private dataset with high precision, and then synthesize a dataset obeying all selected low-dimensional distributions as the synthetic dataset. However, it is difficult to select suitable low-dimensional distributions, which in turn degrades the data utility of resulting synthetic dataset. To improve differentially private dataset synthesis, we propose to train a differentially private AI model for downstream tasks on the original private dataset and utilize the trained model to synthesize datasets. In particular, on the one hand, the AI model satisfies differential privacy so no matter how to use the model does not disclose private information of original private dataset. On the other hand, the AI model is trained to complete the downstream task so the AI model preserves critical information for completing downstream tasks. We utilize the AI model to synthesize datasets to achieve the goal of improving data utility while preserving privacy. Empirical evaluations on four benchmark datasets demonstrate that our proposed DPDSyn consistently outperforms eight state-of-the-art baselines with a maximum improvement of 2.40x in accuracy and 333.73x in synthesis efficiency. Further experiments also validate that DPDSyn has strong scalability across varying data scales.
Problem

Research questions and friction points this paper is trying to address.

differentially private dataset synthesis
data utility
downstream task
privacy-preserving machine learning
synthetic data generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

differentially private synthesis
downstream task guidance
privacy-preserving machine learning
synthetic dataset generation
model-based data utility
🔎 Similar Papers
No similar papers found.