Differentially Private Synthetic Data Generation for Relational Databases

📅 2024-05-29
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
Generating differentially private synthetic data for multi-table relational databases (e.g., academic transcript systems) remains challenging due to structural constraints and utility degradation from flattening. Method: We propose the first end-to-end differentially private (DP) relational synthetic data generation algorithm. It avoids flattening by iteratively calibrating low-order marginals while explicitly enforcing relational constraints—ensuring referential integrity and approximating low-dimensional joint distributions. Contribution/Results: Our approach is the first to decouple arbitrary DP mechanisms from relational structure, achieving both computational efficiency and high-dimensional scalability. It provides rigorous DP guarantees and theoretical utility bounds. Experiments on real-world datasets demonstrate substantial improvements over flattened baselines in query accuracy, cardinality consistency, and JOIN pattern preservation—yielding high-fidelity synthetic relational data.

Technology Category

Machine Learning: PrivacyReasoning under Uncertainty: Relational Probabilistic ModelsData Mining & Knowledge Management: Representing, Reasoning, and Using Provenance, Trust

Application Category

Security and Privacy: Data transparency and provenanceGraph Algorithms and Modeling for the Web: Algorithms and analysis for heterogeneous, signed, attributed, multi-relational, temporal, higher-order, and annotated Web-related graphsUser Modeling, Personalization and Recommendation: User privacy protection in personalized systems
📝 Abstract
Existing differentially private (DP) synthetic data generation mechanisms typically assume a single-source table. In practice, data is often distributed across multiple tables with relationships across tables. In this paper, we introduce the first-of-its-kind algorithm that can be combined with any existing DP mechanisms to generate synthetic relational databases. Our algorithm iteratively refines the relationship between individual synthetic tables to minimize their approximation errors in terms of low-order marginal distributions while maintaining referential integrity. This algorithm eliminates the need to flatten a relational database into a master table (saving space), operates efficiently (saving time), and scales effectively to high-dimensional data. We provide both DP and theoretical utility guarantees for our algorithm. Through numerical experiments on real-world datasets, we demonstrate the effectiveness of our method in preserving fidelity to the original data.
Problem

Research questions and friction points this paper is trying to address.

Privacy Protection
Synthetic Data
Complex Structured Data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic Data Generation
Privacy Preservation
Multi-table Correlation
🔎 Similar Papers
No similar papers found.