Score
Designs and implements algorithms and pipelines that synthesize data records by iteratively selecting and evolving candidate records using progressive selection strategies while enforcing differential privacy. This work includes mechanisms to concentrate and allocate privacy budget on selection steps, perform multi-batch top-1 private selection, and coordinate progressive selection across batches to produce synthetic outputs under strict DP guarantees.
This work addresses the challenge of preserving high-order variable correlations in synthetic tabular data under differential privacy. The authors propose Tab-PE, the first algorithm to extend the Private Evolution framework to tabular data generation. Tab-PE introduces low-overhead heuristic evolutionary operators that efficiently optimize candidate datasets within differential privacy constraints, combined with a private scoring and selection mechanism to simultaneously capture complex correlations and ensure scalability. Experimental results demonstrate that Tab-PE significantly outperforms existing methods on both real-world and synthetic datasets, achieving up to a 10% improvement in classification accuracy over the AIM baseline while running 28 times faster.
This work addresses the challenge of generating high-quality synthetic structured text under strict privacy constraints, where real data are scarce and heavily regulated. Existing differentially private synthesis methods struggle to simultaneously preserve semantic fidelity and structural validity. To overcome this, the authors propose SelPE, a novel framework that concentrates the privacy budget on a multi-batch top-1 progressive selection process and decouples generation into two stages—semantic abstraction followed by structural realization. Candidate samples are evaluated using a multi-channel distance kernel, and diversity is enhanced through a non-private contrastive expansion mechanism that incurs no additional privacy cost. Experiments demonstrate that, under stringent differential privacy guarantees and limited sample sizes, SelPE significantly improves the structural validity, semantic fidelity, and downstream utility of the synthesized data.
Generating differentially private synthetic data for multi-table relational databases (e.g., academic transcript systems) remains challenging due to structural constraints and utility degradation from flattening. Method: We propose the first end-to-end differentially private (DP) relational synthetic data generation algorithm. It avoids flattening by iteratively calibrating low-order marginals while explicitly enforcing relational constraints—ensuring referential integrity and approximating low-dimensional joint distributions. Contribution/Results: Our approach is the first to decouple arbitrary DP mechanisms from relational structure, achieving both computational efficiency and high-dimensional scalability. It provides rigorous DP guarantees and theoretical utility bounds. Experiments on real-world datasets demonstrate substantial improvements over flattened baselines in query accuracy, cardinality consistency, and JOIN pattern preservation—yielding high-fidelity synthetic relational data.
Existing differential privacy (DP) synthetic data methods suffer substantial utility degradation when input data contain missing values. Method: This paper formally defines the DP synthetic data generation problem under missingness, characterizing how missingness mechanisms—Missing Completely at Random (MCAR), Missing at Random (MAR), and Missing Not at Random (MNAR)—propagate privacy loss and affect theoretical privacy bounds. Building on this analysis, we propose three adaptive synthesis strategies, each tailored to a specific missingness mechanism and privacy budget. Contribution/Results: Our key innovation is integrating missingness mechanism modeling directly into the DP synthetic data framework, enabling joint optimization of privacy protection and data utility. Evaluated on four real-world datasets with natural missingness, our approach achieves significant utility gains—improving downstream machine learning task accuracy by an average of 8.3%—while yielding tighter analytical privacy upper bounds compared to baseline methods.
Existing tabular data synthesis methods lack a unified, comparable evaluation framework due to fragmented metrics and missing standardized benchmarks—particularly concerning privacy guarantees (differential privacy vs. heuristic approaches) and model paradigms (diffusion models, LLMs vs. statistical methods). Method: We introduce the first systematic evaluation framework for privacy-preserving tabular synthesis, featuring a three-dimensional quantitative metric system—fidelity, privacy, and utility—and a differentiable, unified objective function enabling fair cross-paradigm comparison across diffusion models, LLM-based synthesizers, and marginal-distribution methods. The framework integrates formal differential privacy verification, multi-scale statistical utility assessment (e.g., MMD, JS divergence), adversarial privacy attack benchmarks, and downstream task generalization tests. Results: Extensive experiments across 12 real-world datasets and 8 synthesizers reveal fundamental performance boundaries and trade-off patterns, providing empirical guidance and concrete improvement pathways for next-generation privacy-enhanced synthetic data generation.
This work addresses the problem of efficiently generating synthetic data under differential privacy for a given family of queries. By parameterizing the problem with the treewidth of the query family’s associated graph, the authors establish—for the first time—that the problem is fixed-parameter tractable. They propose a unified dynamic programming framework that integrates linear programming duality-based separation, subsampled private multiplicative weights, and Gibbs sampling techniques. This approach achieves theoretically optimal error rates across the full parameter regime, significantly enhancing both the scalability and practical utility of differentially private synthetic data generation.
This work addresses the vulnerability of differentially private synthetic data in protecting anomalous individuals—such as patients with rare diseases—who face significantly higher success rates in membership inference attacks. To mitigate this risk, the authors propose a risk-balanced differentially private synthesis framework that first evaluates the anomaly level of each record using a small privacy budget and then inversely weights records by their risk during generative model training to attenuate the influence of high-risk samples. This approach provides stronger privacy guarantees by incorporating record-level risk awareness into differentially private synthesis for the first time, enabling targeted protection for highly anomalous individuals and yielding a closed-form per-record privacy bound. Experiments on both synthetic and real-world datasets (e.g., Breast Cancer, Adult) demonstrate a substantial reduction in membership inference success against high-anomaly records, with performance contingent on the synergy between the risk scorer and the synthesis pipeline.
This work addresses fairness concerns in differentially private synthetic data, which often arise from spurious associations between sensitive attributes and outcomes. To mitigate this issue, the authors propose PrivCI, a method that incorporates conditional independence (CI) constraints into the differentially private data synthesis process to eliminate such biases. The core innovation lies in a CI-aware greedy minimum spanning tree algorithm that integrates feasibility checks and the exponential mechanism during Kruskal’s construction, enabling graph structure learning that jointly preserves privacy, fairness, and data fidelity. Experimental results demonstrate that PrivCI significantly outperforms existing approaches in data fidelity and predictive accuracy on standard fairness benchmarks while strictly adhering to prescribed CI constraints.
This study addresses the instability and noise sensitivity of greedy construction in differentially private prompt optimization under strict privacy budgets. To overcome these limitations, we propose DP-ES, a method that replaces word-by-word greedy search with a population-based evolutionary strategy. By maintaining a complete population of prompts for mutation and evaluation, DP-ES strictly confines privacy consumption to the Gaussian-sampled evaluation phase and incorporates Gumbel-smoothed selection to achieve efficient and secure prompt optimization. Experimental results demonstrate that DP-ES attains 88.1% accuracy on GSM8K, representing a 38.6 percentage point improvement over baselines. Furthermore, it reduces variance by a factor of nine, accelerates inference speed by 2.5 times, and requires fewer privacy queries, highlighting its effectiveness for privacy-preserving prompt engineering.
为解决隐私保护下的表格数据共享问题,提出TabSSD方法,利用大语言模型设计合成策略而非直接生成记录,平衡了统计准确性、预测效用和隐私风险。