🤖 AI Summary
Generating differentially private synthetic data for multi-table relational databases (e.g., academic transcript systems) remains challenging due to structural constraints and utility degradation from flattening.
Method: We propose the first end-to-end differentially private (DP) relational synthetic data generation algorithm. It avoids flattening by iteratively calibrating low-order marginals while explicitly enforcing relational constraints—ensuring referential integrity and approximating low-dimensional joint distributions.
Contribution/Results: Our approach is the first to decouple arbitrary DP mechanisms from relational structure, achieving both computational efficiency and high-dimensional scalability. It provides rigorous DP guarantees and theoretical utility bounds. Experiments on real-world datasets demonstrate substantial improvements over flattened baselines in query accuracy, cardinality consistency, and JOIN pattern preservation—yielding high-fidelity synthetic relational data.
📝 Abstract
Existing differentially private (DP) synthetic data generation mechanisms typically assume a single-source table. In practice, data is often distributed across multiple tables with relationships across tables. In this paper, we introduce the first-of-its-kind algorithm that can be combined with any existing DP mechanisms to generate synthetic relational databases. Our algorithm iteratively refines the relationship between individual synthetic tables to minimize their approximation errors in terms of low-order marginal distributions while maintaining referential integrity. This algorithm eliminates the need to flatten a relational database into a master table (saving space), operates efficiently (saving time), and scales effectively to high-dimensional data. We provide both DP and theoretical utility guarantees for our algorithm. Through numerical experiments on real-world datasets, we demonstrate the effectiveness of our method in preserving fidelity to the original data.